Apparatus and method for controlling and debugging hardware lockstep mismatch
By designing multiple processing elements in the processor to execute instructions in redundant mode, and using comparator and masking circuit to process fault indications, the problem of inefficient hardware lock step error handling in the prior art is solved, and more efficient fault processing and debugging is achieved.
Patent Information
- Application Number
- CN202411499251.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-28
- Filing Date
- 2024-10-25
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is inefficient when dealing with hardware lock step errors, especially mismatch errors, which leads to frequent shutdowns and debugging difficulties.
A processor is designed, including multiple processing elements that execute the same instructions in redundant processing mode, generate a result signal, and compare the result signal through a comparator circuit to generate a fault indication. The masking circuit masks certain minor or expected failures through the setting of the mask matrix, avoiding unnecessary shutdowns.
Improves the efficiency of hardware lock step error handling, reduces unnecessary shutdowns, simplifies the debugging process, accurately identify the source of failures, and achieves more efficient operation and shorter downtime.
Smart Images

Figure CN120066880A_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to the field of computer processors. More specifically, the present invention relates to an apparatus and method for controlling and debugging hardware lockstep errors such as mis-compare. Background Art
[0002] Hardware Lockstep (HWLS) is a feature in the server industry that supports Reliability and Security (RAS). Hardware Lockstep runs two entities in the system in an active-shadow configuration, which means that the two entities in the module are in a state where they run or perform the same operations in a cycle-by-cycle modular manner.
[0003] In a typical implementation, a comparator tree compares "critical feature" signals from two entities in each cycle, and any divergence observed in this traffic results in a so-called "mis-match". Subsequently, a fault signal is issued, and then a hard machine check error is recorded in the machine check library. This divergence is caused by a defect in one of the entities.
[0004] In the current implementation, the system actions for issuing these types of fault detections are inefficient. In most cases, a halt state is triggered via a micro breakpoint, and traditional debugging methods are used to try to find the root cause. Any non-deterministic but expected divergence also results in lockstep differences, which in turn lead to frequent shutdowns of the client. Summary of the Invention
[0005] According to one aspect of the present application, a processor is provided, including: a plurality of processing elements operable in a redundant processing mode, each of the plurality of processing elements for executing the same plurality of instructions and generating a corresponding plurality of result signals; a comparator circuit for comparing corresponding result signals among the corresponding plurality of result signals, the comparator circuit for: generating one or more fault indications when a first one or more result signals generated by a first processing element are different from corresponding second one or more result signals generated by a second processing element; and a masking circuit for masking a first fault indication among the one or more fault indications when a corresponding mask bit in a mask matrix is set to a first value.
[0006] In another aspect of the present application, a method is provided, including: operating a plurality of processing elements in a redundancy processing mode, each of the plurality of processing elements being configured to execute the same plurality of instructions and generate corresponding plurality of result signals; comparing the corresponding result signals among the corresponding plurality of result signals, and generating one or more fault indications when the first one or more result signals generated by a first processing element are different from the corresponding second one or more result signals generated by a second processing element; and masking a first fault indication among the one or more fault indications when a corresponding mask bit of a mask matrix is set to a first value.
[0007] In yet another aspect of the present application, a machine-readable medium is provided, having program code stored thereon, which when executed by a machine, causes the machine to perform operations, the operations including: operating a plurality of processing elements in a redundancy processing mode, each of the plurality of processing elements executing the same plurality of instructions and generating corresponding plurality of result signals; comparing the corresponding result signals among the corresponding plurality of result signals, and generating one or more fault indications when the first one or more result signals generated by a first processing element are different from the corresponding second one or more result signals generated by a second processing element; and masking a first fault indication among the one or more fault indications when a corresponding mask bit of a mask matrix is set to a first value. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] A better understanding of the present invention can be obtained from the following detailed description in conjunction with the accompanying drawings, in which:
[0009] Figure 1 An example computer system architecture is illustrated;
[0010] Figure 2 A processor including a plurality of cores is illustrated;
[0011] Figure 3A A plurality of stages of a processing pipeline are illustrated;
[0012] Figure 3B Details of an embodiment of a core are illustrated;
[0013] Figure 4 An execution circuit according to an embodiment is illustrated;
[0014] Figure 5 An embodiment of a register architecture is illustrated;
[0015] Figure 6 An example of an instruction format is illustrated;
[0016] Figure 7 An addressing technique according to an embodiment is illustrated;
[0017] Figure 8 Illustrates an embodiment of an instruction prefix;
[0018] Figures 9A - 9D Illustrates an embodiment of how to use the R, X, and B fields of the prefix;
[0019] Figures 10A - 10B Illustrates an example of a second instruction prefix;
[0020] Figure 11 Illustrates the payload bytes of an embodiment of an instruction prefix;
[0021] Figure 12 Illustrates instruction conversion and binary translation implementations;
[0022] Figure 13 Illustrates a comparator circuit and a masking circuit for processing redundant execution results according to an embodiment of the present invention;
[0023] Figure 14 Illustrates a set of per-group results according to an embodiment;
[0024] Figure 15 Illustrates a method according to an embodiment of the present invention; and
[0025] Figure 16 Illustrates an example of dual-core lockstep operation according to an embodiment of the present invention. Detailed Description
[0026] In the following description, for purposes of illustration, numerous specific details are set forth in order to provide a thorough understanding of embodiments of the invention described below. However, those skilled in the art will appreciate that embodiments of the invention may be practiced without some of these specific details. In other instances, well-known structures and devices are shown in block diagram form to avoid obscuring the underlying principles of embodiments of the invention.
[0027] Exemplary Computer Architecture
[0028] The exemplary computer architecture is described in detail below. Other system designs and configurations known in the art for laptops, desktops, handheld PCs, personal digital assistants, engineering workstations, servers, network devices, network hubs, switches, embedded processors, digital signal processors (DSPs), graphics devices, video game devices, set-top boxes, microcontrollers, cellular phones, portable media players, handheld devices, and various other electronic devices are also suitable. In general, a wide variety of systems or electronic devices capable of incorporating the processors and / or other execution logic disclosed herein are generally suitable.
[0029] Figure 1 An embodiment of an exemplary system is illustrated. The multi-processor system 100 is a point-to-point interconnect system and includes multiple processors, including a first processor 170 and a second processor 180 coupled via a point-to-point interconnect 150. In some embodiments, the first processor 170 and the second processor 180 are homogeneous. In some embodiments, the first processor 170 and the second processor 180 are heterogeneous.
[0030] Processors 170 and 180 are shown as including integrated memory controller (IMC) unit circuits 172 and 182, respectively. Processor 170 also includes point-to-point (P-P) interfaces 176 and 178 as part of its interconnect controller unit; similarly, the second processor 180 includes P-P interfaces 186 and 188. Processors 170, 180 can exchange information via the P-P interfaces 150 using the P-P interface circuits 178, 188. The IMCs 172 and 182 couple the processors 170, 180 to their respective memories, namely memory 132 and memory 134, which may be part of the main memory locally attached to each processor.
[0031] Processors 170, 180 can each exchange information with the chipset 190 via individual P-P interconnects 152, 154 using the point-to-point interface circuits 176, 194, 186, 198. The chipset 190 can optionally exchange information with the coprocessor 138 via the high-performance interface 192. In some embodiments, the coprocessor 138 is a dedicated processor, e.g., a high-throughput MIC processor, a network or communication processor, a compression engine, a graphics processor, a GPGPU, an embedded processor, etc.
[0032] A shared cache (not shown) may be included in either processor 170, 180, or outside both processors, but connected to the processors via a P-P interconnect, such that local cache information of either or both processors can be stored in the shared cache when the processors are placed in a low-power mode.
[0033] The chipset 190 can be coupled to the first interconnect 116 via an interface 196. In some embodiments, the first interconnect 116 can be a Peripheral Component Interconnect (PCI) interconnect, or an interconnect such as a PCI Express interconnect or another I / O interconnect. In some embodiments, one of the interconnects is coupled to a power control unit (PCU) 117, which can include circuitry, software, and / or firmware to perform power management operations regarding the processors 170, 180, and / or the coprocessor 138. The PCU 117 provides control information to the voltage regulators such that the voltage regulators generate appropriate regulated voltages. The PCU 117 also provides control information to control the generated operating voltages. In various embodiments, the PCU 117 can include multiple power management logic units (circuits) to perform hardware-based power management. Such power management can be fully processor-controlled (e.g., controlled by various processor hardware and can be triggered by workload and / or power constraints, thermal constraints, or other processor constraints), and / or the power management can be performed in response to an external source (e.g., a platform or a power management source or system software).
[0034] The PCU 117 is illustrated as existing as a separate logic from the processor 170 and / or the processor 180. In other cases, the PCU 117 can execute on a given one or more cores in the core (not shown) of the processor 170 or 180. In some cases, the PCU 117 can be implemented as a microcontroller (dedicated or general-purpose) or other control logic, which is configured to execute its own dedicated power management code (sometimes referred to as P-code). In still other embodiments, the power management operations to be performed by the PCU 117 can be implemented external to the processor, for example, by a separate power management integrated circuit (PMIC) or another component external to the processor. In still other embodiments, the power management operations to be performed by the PCU 117 can be implemented within the BIOS or other system software.
[0035] Various I / O devices 114 may be coupled to the first interconnect 116, and an interconnect (bus) bridge 118 that couples the first interconnect 116 to the second interconnect 120. In some embodiments, one or more additional processors 115, such as a coprocessor, a high throughput MIC processor, a GPGPU, an accelerator (e.g., a graphics accelerator or a digital signal processing (DSP) unit), a field programmable gate array (FPGA), or any other processor, are coupled to the first interconnect 116. In some embodiments, the second interconnect 120 may be a low pin count (LPC) interconnect. Various devices may be coupled to the second interconnect 120, including, for example, a keyboard and / or mouse 122, a communication device 127, and a storage unit circuit 128. The storage unit circuit 128 may be a disk drive or other mass storage device, and in some embodiments, it may include instructions / code and data 130. Additionally, audio I / O 124 may be coupled to the second interconnect 120. Note that other architectures are possible in addition to the point-to-point architecture described above. For example, a system such as the multiprocessor system 100 may implement a multi-drop interconnect or other such architecture instead of a point-to-point architecture.
[0036] Exemplary Core Architectures, Processors, and Computer Architectures
[0037] Processor cores may be implemented in different ways, for different purposes, and in different processors. For example, implementations of these cores may include: 1) general purpose in-order cores, for general purpose computing purposes; 2) high performance general purpose out-of-order cores for general purpose computing purposes; 3) special purpose cores, primarily for graphics and / or scientific (throughput) computing purposes. Implementations of different processors may include: 1) a CPU, including one or more general purpose in-order cores for general purpose computing purposes and / or one or more general purpose out-of-order cores for general purpose computing purposes; and 2) a coprocessor, including one or more special purpose cores primarily for graphics and / or scientific (throughput) purposes. Such different processors lead to different computer system architectures, which may include: 1) coprocessors on separate chips from the CPU; 2) coprocessors on separate dies in the same package as the CPU; 3) coprocessors on the same die as the CPU (in which case such coprocessors are sometimes referred to as dedicated logic, such as integrated graphics and / or scientific (throughput) logic, or as dedicated cores); and 4) systems on a chip, which may include the above-described coprocessors as well as additional functionality on the same die as the described CPU (sometimes referred to as application core(s) or application processor(s). An exemplary core architecture is described next, followed by a description of an exemplary processor and computer architecture.
[0038] Figure 2 A block diagram of an embodiment of a processor 200 is illustrated, which may have more than one core, may have an integrated memory controller, and may have integrated graphics. The solid line box illustrates the processor 200 with a single core 202A, a system agent 210, and a set of one or more interconnected controller unit circuits 216, while the optionally added dashed line box illustrates an alternative processor 200 with multiple cores 202(A)-(N), a set of one or more integrated memory control unit circuits 214 and dedicated logic 208 in the system agent unit circuit 210, and a set of one or more interconnected controller unit circuits 216. Note that the processor 200 may be Figure 1 One of the processors 170 or 180 or the coprocessors 138 or 115.
[0039] Accordingly, different implementations of the processor 200 may include: 1) a CPU, where the dedicated logic 208 is integrated graphics and / or scientific (throughput) logic (which may include one or more cores, not shown), and the cores 202(A)-(N) are one or more general-purpose cores (e.g., general-purpose in-order cores, general-purpose out-of-order cores, or a combination of both); 2) a coprocessor, where the cores 202(A)-(N) are a large number of dedicated cores mainly for graphics and / or scientific (throughput) purposes; and 3) a coprocessor, where the cores 202(A)-(N) are a large number of general-purpose in-order cores. Accordingly, the processor 200 may be a general-purpose processor, a coprocessor, or a special-purpose processor, e.g., a network or communication processor, a compression engine, a graphics processor, a GPGPU (general-purpose graphics processing unit circuit), a high-throughput many integrated core (MIC) coprocessor (including 30 or more cores), an embedded processor, etc. The processor may be implemented on one or more chips. The processor 200 may be part of one or more substrates and / or may be implemented on one or more substrates using any of several process technologies, such as BiCMOS, CMOS, or NMOS.
[0040] The memory hierarchy includes one or more levels of cache unit circuits 204(A)-(N) within the cores 202(A)-(N), a group of or one or more shared cache unit circuits 206, and external memory (not shown) coupled to the group of integrated memory controller unit circuits 214. The group of one or more shared cache unit circuits 206 may include one or more intermediate-level caches (e.g., level 2 (L2), level 3 (L3), level 4 (L4), or other levels of cache), such as a last level cache (LLC), and / or combinations thereof. Although in some embodiments a ring-based interconnect network circuit 212 interconnects the dedicated logic 208 (e.g., integrated graphics logic), the group of shared cache unit circuits 206, and the system agent unit circuit 210, alternative embodiments use any number of well-known techniques to interconnect such units. In some embodiments, coherence is maintained between one or more of the shared cache unit circuits 206 and the cores 202(A)-(N).
[0041] In some embodiments, one or more of cores 202(A)-(N) are capable of multithreading. The system agent unit circuitry 210 includes those components that coordinate and operate cores 202(A)-(N). The system agent unit circuitry 210 may include, for example, power control unit (PCU) circuitry and / or display unit circuitry (not shown). The PCU may be or may include the logic and components required to regulate the power states of cores 202(A)-(N) and / or dedicated logic 208 (e.g., integrated graphics logic). The display unit circuitry is used to drive one or more externally connected displays.
[0042] Cores 202(A)-(N) may be homogeneous or heterogeneous with respect to the architectural instruction set; that is, two or more of cores 202(A)-(N) may be capable of executing the same instruction set, while other cores may be capable of executing only a subset of that instruction set or a different instruction set.
[0043] Exemplary Core Architectures
[0044] In-order and Out-of-order Core Block Diagrams
[0045] The block diagrams of FIG. 3(A) illustrate both an exemplary in-order pipeline and an exemplary register renaming, out-of-order issue / execution pipeline in accordance with an embodiment of the present invention. The block diagrams of FIG. 3(B) illustrate both an exemplary embodiment of an in-order architecture core to be included in a processor and an exemplary register renaming, out-of-order issue / execution architecture core in accordance with an embodiment of the present invention. The solid block diagrams in FIGS. 3(A)-3(B) illustrate the in-order pipeline and in-order core, while the optionally added dashed block diagrams illustrate the register renaming, out-of-order issue / execution pipeline and core. Given that the in-order aspect is a subset of the out-of-order aspect, the out-of-order aspect will be described.
[0046] In FIG. 3(A), the processor pipeline 300 includes a fetch stage 302, an optional length decoding stage 304, a decode stage 306, an optional allocation stage 308, an optional rename stage 310, a schedule (also known as dispatch or issue) stage 312, an optional register read / memory read stage 314, an execute stage 316, a writeback / memory write stage 318, an optional exception handling stage 322, and an optional commit stage 324. One or more operations may be performed in each of these processor pipeline stages. For example, during the fetch stage 302, one or more instructions are fetched from an instruction memory. During the decode stage 306, the fetched one or more instructions may be decoded, an address using a forwarding register port (e.g., a load store unit (LSU) address) may be generated, and branch forwarding (e.g., immediate offset or link register (LR)) may be performed. In one embodiment, the decode stage 306 and the register read / memory read stage 314 may be combined into one pipeline stage. In one embodiment, during the execute stage 316, the decoded instructions may be executed, LSU address / data pipelining to an Advanced Microcontroller Bus (AHB) interface may be performed, multiplication and addition operations may be performed, arithmetic operations with branch results may be performed, and so on.
[0047] As an example, an exemplary register renaming, out-of-order issue / execution core architecture may implement the pipeline 300 as follows: 1) Instruction fetch 338 performs the fetch and length decoding stages 302 and 304; 2) The decode unit circuit 340 performs the decode stage 306; 3) The rename / allocator unit circuit 352 performs the allocation stage 308 and the rename stage 310; 4) The scheduler unit circuit 356 performs the schedule stage 312; 5) The physical register file unit circuit 358 and the memory unit circuit 370 perform the register read / memory read stage 314; The execution cluster 360 performs the execute stage 316; 6) The memory unit circuit 370 and the physical register file unit circuit 358 perform the writeback / memory write stage 318; 7) Various units (unit circuits) may be involved in the exception handling stage 322; and 8) The retirement unit circuit 354 and the physical register file unit circuit 358 perform the commit stage 324.
[0048] FIG. 3(B) shows that the processor core 390 includes a front-end unit circuit 330 coupled to an execution engine unit circuit 350, and both are coupled to a memory unit circuit 370. The core 390 can be a reduced instruction set computing (RISC) core, a complex instruction set computing (CISC) core, a very long instruction word (VLIW) core, or a hybrid or alternative core type. As another option, the core 390 can be a specialized core, such as, for example, a network or communication core, a compression engine, a co-processor core, a general purpose computing graphics processing unit (GPGPU) core, a graphics core, and so on.
[0049] The front-end unit circuit 330 may include a branch prediction unit circuit 332, which is coupled to an instruction cache unit circuit 334, which is coupled to an instruction translation lookaside buffer (TLB) 336, which is coupled to an instruction fetch unit circuit 338, which is coupled to a decode unit circuit 340. In one embodiment, the instruction cache unit circuit 334 is included in the memory unit circuit 370 rather than the front-end unit circuit 330. The decode unit circuit 340 (or decoder) may decode the instructions and generate one or more micro-operations, micro-code entry points, micro-instructions, other instructions, or other control signals as output, which are decoded from the original instructions, or otherwise reflect the original instructions, or are derived from the original instructions. The decode unit circuit 340 may also include an address generation unit circuit (AGU, not shown). In one embodiment, the AGU uses a forwarded register port to generate LSU addresses and may further perform branch forwarding (e.g., immediate offset branch forwarding, LR register branch forwarding, etc.). The decode unit circuit 340 may be implemented using a variety of different mechanisms. Examples of suitable mechanisms include, but are not limited to, lookup tables, hardware implementations, programmable logic arrays (PLAs), micro-code read only memories (ROMs), etc. In one embodiment, the core 390 includes a micro-code ROM (not shown) or other medium that stores micro-code for certain macro instructions (e.g., in the decode unit circuit 340 or otherwise within the front-end unit circuit 330). In one embodiment, the decode unit circuit 340 includes a micro-operation (micro-op) or operation cache (not shown) to hold / cache the decoded operations, micro-tags, or micro-operations generated during decoding or other stages of the processor pipeline 300. The decode unit circuit 340 may be coupled to a rename / allocator unit circuit 352 in the execution engine unit circuit 350.
[0050] The execution engine circuit 350 includes a rename / allocator unit circuit 352, which is coupled to a retirement unit circuit 354 and a set of one or more scheduler circuits 356. The scheduler circuits 356 represent any number of different schedulers, including reservation stations, a central instruction window, and so on. In some embodiments, the scheduler circuits 356 may include an arithmetic logic unit (ALU) scheduler / scheduling circuit, an ALU queue, an arithmetic generation unit (AGU) scheduler / scheduling circuit, an AGU queue, and so on. The scheduler circuits 356 are coupled to a physical register file circuit 358. Each of the physical register file circuits 358 represents one or more physical register files, and different physical register files among these physical register files store one or more different data types, such as scalar integers, scalar floating points, packed integers, packed floating points, vector integers, vector floating points, status (e.g., an instruction pointer that is the address of the next instruction to be executed), and so on. In one embodiment, the physical register file unit circuit 358 includes a vector register unit circuit, a write mask register unit circuit, and a scalar register unit circuit. These register units can provide architected vector registers, vector mask registers, general-purpose registers, and so on. The physical register file unit circuit 358 overlaps with the retirement unit circuit 354 (also referred to as a retirement queue or a commit queue) to illustrate various ways that can be used to implement register renaming and out-of-order execution (e.g., using one or more reorder buffers (ROBs) and one or more retirement register files; using one or more future heaps, one or more history buffers, and one or more retirement register files; using a register map and a pool of registers; and so on). The retirement unit circuit 354 and the one or more physical register file circuits 358 are coupled to the one or more execution clusters 360. The one or more execution clusters 360 include a set of one or more execution unit circuits 362 and a set of one or more memory access circuits 364. The execution unit circuits 362 can perform various arithmetic, logical, floating-point, or other types of operations (e.g., shifts, additions, subtractions, multiplications) on various types of data (e.g., scalar floating points, packed integers, packed floating points, vector integers, vector floating points). Although some embodiments may include several execution units or execution unit circuits dedicated to a specific function or set of functions, other embodiments may include only one execution unit circuit or multiple execution units / execution unit circuits that perform all functions.The scheduler circuit 356, the physical register file unit circuit 358, and the execution cluster(s) 360 are shown as potentially multiple because some embodiments create separate pipelines for certain types of data / operations (e.g., scalar integer pipelines, scalar floating point / tight integer / tight floating point / vector integer / vector floating point pipelines, and / or memory access pipelines, each having its own scheduler circuit, physical register file unit circuit, and / or execution cluster—and in the case of a separate memory access pipeline, some embodiments are implemented where only the execution cluster of this pipeline has the memory access unit circuit 364). It should also be understood that in the case of using separate pipelines, one or more of these pipelines may be out-of-order issue / execution while the rest are in-order.
[0051] In some embodiments, the execution engine unit circuit 350 may perform load / store unit (LSU) address / data pipelining to an advanced microcontroller bus (AHB) interface (not shown), as well as address phase and write-back, data phase load, store, and branch.
[0052] The set of memory access circuits 364 is coupled to the memory unit circuit 370, which includes a data TLB unit circuit 372 that is coupled to a data cache circuit 374, which is coupled to a level 2 (L2) cache circuit 376. In one exemplary embodiment, the memory access unit circuit 364 may include a load unit circuit, a store address unit circuit, and a store data unit circuit, each of which is coupled to the data TLB circuit 372 in the memory unit circuit 370. The instruction cache circuit 334 is further coupled to the level 2 (L2) cache unit circuit 376 in the memory unit circuit 370. In one embodiment, the instruction cache 334 and the data cache 374 are combined into a single instruction and data cache (not shown) in the L2 cache unit circuit 376, a level 3 (L3) cache unit circuit (not shown), and / or main memory. The L2 cache unit circuit 376 is coupled to one or more other levels of cache and ultimately to main memory.
[0053] The core 390 may support one or more instruction sets (e.g., the x86 instruction set (with some extensions added with newer versions); the MIPS instruction set; the ARM instruction set (with optional additional extensions such as NEON)), including the instruction(s) described herein. In one embodiment, the core 390 includes logic to support packed data instruction set extensions (e.g., AVX1, AVX2), allowing operations used by many multimedia applications to be performed using packed data.
[0054] Exemplary execution unit circuit
[0055] Figure 4 An embodiment of an execution unit circuit is illustrated, such as the execution unit circuit 362 of FIG. 3(B). As shown, the execution unit circuit 362 may include one or more ALU circuits 401, vector / SIMD unit circuits 403, load / store unit circuits 405, and / or branch / jump unit circuits 407. The ALU circuit 401 performs integer arithmetic and / or Boolean operations. The vector / SIMD unit circuit 403 performs vector / SIMD operations on packed data (e.g., SIMD / vector registers). The load / store unit circuit 405 executes load and store instructions to load data from memory into registers or store data from registers to memory. The load / store unit circuit 405 may also generate addresses. The branch / jump unit circuit 407 causes a branch or jump to a certain memory address depending on the instruction. The floating-point unit (FPU) circuit 409 performs floating-point arithmetic. The width of the execution unit circuit 362 varies depending on the embodiment and may range from 16 bits to 1024 bits. In some embodiments, two or more smaller execution units are logically combined to form a larger execution unit (e.g., two 128-bit execution units are logically combined to form a 256-bit execution unit).
[0056] Exemplary register architecture
[0057] Figure 5 is a block diagram of a register architecture 500 according to some embodiments. As shown, there are vector / SIMD registers 510, whose widths vary from 128 bits to 1024 bits. In some embodiments, the vector / SIMD registers 510 are physically 512 bits, and depending on the mapping, only some of the lower bits are used. For example, in some embodiments, the vector / SIMD register 510 is a 512-bit ZMM register: the lower 256 bits are used for YMM registers, and the lower 128 bits are used for XMM registers. Thus, there is register coverage. In some embodiments, the vector length field selects between a maximum length and one or more other shorter lengths, where each such shorter length is half the length of the previous length; scalar operations are operations performed on the lowest-order data element positions in the ZMM / YMM / XMM registers; the higher-order data element positions either remain the same as they were before the instruction or are zeroed, depending on the embodiment.
[0058] In some embodiments, the register architecture 500 includes write mask / predicate registers 515. For example, in some embodiments, there are eight write mask / predicate registers (sometimes referred to as k0 through k7), each of which is sized 16 bits, 32 bits, 64 bits, or 128 bits. The write mask / predicate registers 515 can permit coalescing (e.g., permit any subset of elements in a destination to be immune from update during the execution of any operation) and / or zeroing (e.g., a zeroing vector mask permits any subset of elements in a destination to be zeroed during the execution of any operation). In some embodiments, each data element position in a given write mask / predicate register 515 corresponds to a data element position in a destination. In other embodiments, the write mask / predicate registers 515 are scalable and consist of a set number of enable bits for a given vector element (e.g., eight enable bits for each 64-bit vector element).
[0059] The register architecture 500 includes a plurality of general-purpose registers 525. These registers can be 16 bits, 32 bits, 64 bits, etc., and are capable of being used for scalar operations. In some embodiments, these registers are referenced by the names RAX, RBX, RCX, RDX, RBP, RSI, RDI, RSP, and R8 through R15.
[0060] In some embodiments, the register architecture 500 includes scalar floating-point registers 545, which are used for scalar floating-point operations on 32 / 64 / 80-bit floating-point data using the x87 instruction set extensions, or as MMX registers to perform operations on 64-bit packed integer data, and to hold operands for some operations performed between MMX and XMM registers.
[0061] One or more flag registers 540 (e.g., EFLAGS, RFLAGS, etc.) store status and control information for arithmetic, comparison, and system operations. For example, one or more flag registers 540 can store condition code information such as carry, parity, auxiliary carry, zero, sign, and overflow. In some embodiments, one or more flag registers 540 are referred to as program status and control registers.
[0062] The segment registers 520 contain segment points for accessing memory. In some embodiments, these registers are referred to by the names CS, DS, SS, ES, FS, and GS.
[0063] A machine-specific register (MSR) 535 controls and reports processor performance. Most MSRs 535 handle system-related functions and are not accessible to applications. The machine check register 560 consists of control, status, and error-reporting MSRs for detecting and reporting hardware errors.
[0064] One or more instruction pointer registers 530 store instruction pointer values. One or more control registers 555 (e.g., CR0 - CR4) determine the operating mode of the processor (e.g., processors 170, 180, 138, 115, and / or 200) and the characteristics of the currently executing task. Debug registers 550 control and allow monitoring of the debugging operations of the processor or core.
[0065] Memory management registers 565 specify the locations for data structures in protected mode memory management. These registers can include the GDTR, IDRT, task registers, and LDTR registers.
[0066] Alternative embodiments of the present invention can use wider or narrower registers. Additionally, alternative embodiments of the present invention can use more, fewer, or different register banks and registers.
[0067] Instruction set
[0068] An instruction set architecture (ISA) can include one or more instruction formats. A given instruction format can define various fields (e.g., number of bits, position of bits) to specify the operation to be performed (e.g., opcode) and the operand(s) on which the operation is to be performed and / or other data field(s) (e.g., mask), etc. Some instruction formats are further decomposed by the definition of instruction templates (or sub-formats). For example, an instruction template of a given instruction format can be defined as having different subsets of the fields of that instruction format (the included fields are typically in the same order, but at least some have different bit positions as fewer fields are included) and / or can be defined as having a given field interpreted in a different way. Thus, each instruction of the ISA is expressed using a given instruction format (and if defined, using a given instruction template within that instruction format's instruction templates), and includes fields for specifying the operation and operands. For example, an exemplary ADD instruction has a specific opcode and an instruction format that includes an opcode field to specify the opcode and an operand field to select the operands (source 1 / destination and source 2); and the occurrence of this ADD instruction in the instruction stream will have specific contents in the operand fields that select the specific operands.
[0069] Exemplary Instruction Formats
[0070] Embodiments of the (one or more) instructions described herein may be implemented in different formats. Additionally, exemplary systems, architectures, and pipelines are detailed below. Embodiments of the (one or more) instructions may be executed on such systems, architectures, and pipelines, but are not limited to those detailed.
[0071] Figure 6 An embodiment of an instruction format is illustrated. As shown, an instruction may include multiple components, including but not limited to one or more fields for: one or more prefixes 601, an opcode 603, addressing information 605 (e.g., register identifiers, memory addressing information, etc.), a displacement value 607, and / or an immediate value 609. Note that some instructions utilize some or all of the fields of this format, while other instructions may use only the fields of the opcode 603. In some embodiments, the illustrated order is the order in which these fields are to be encoded; however, it should be understood that in other embodiments, these fields may be encoded in a different order, combined, etc.
[0072] The (one or more) prefix fields 601, when used, modify the instruction. In some embodiments, one or more prefixes are used for repeat string instructions (e.g., 0xF0, 0xF2, 0xF3, etc.) to provide segment overrides (e.g., 0x2E, 0x36, 0x3E, 0x26, 0x64, 0x65, 0x2E, 0x3E, etc.), to perform bus lock operations, and / or to change operand (e.g., 0x66) and address size (e.g., 0x67). Certain instructions require mandatory prefixes (e.g., 0x66, 0xF2, 0xF3, etc.). Some of these prefixes may be considered "traditional" prefixes. Other prefixes (one or more examples of which are detailed herein) indicate and / or provide further capabilities, such as specifying particular registers, etc. These other prefixes typically follow the "traditional" prefixes.
[0073] The opcode field 603 is used to at least partially define the operation to be performed upon decoding of the instruction. In some embodiments, the length of the primary opcode encoded in the opcode field 603 is 1, 2, or 3 bytes. In other embodiments, the primary opcode may be of a different length. An additional 3-bit opcode field is sometimes encoded in another field.
[0074] The addressing field 605 is used to address one or more operands of the instruction, such as a location in memory or one or more registers. Figure 7Illustrates an embodiment of the addressing field 605. In this illustration, an optional ModR / M byte 702 and an optional Scale, Index, Base (SIB) byte 704 are shown. The ModR / M byte 702 and the SIB byte 704 are used to encode up to two operands of an instruction, where each operand is either a direct register or an effective memory address. Note that each of these fields is optional, i.e., not all instructions include one or more of these fields. The MOD R / M byte 702 includes a MOD field 742, a register field 744, and an R / M field 746.
[0075] The content of the MOD field 742 differentiates between memory access and non-memory access modes. In some embodiments, when the value of the MOD field 742 is b11, the register direct addressing mode is utilized; otherwise, register indirect addressing is used.
[0076] The register field 744 can encode a destination register operand or a source register operand, or can encode an opcode extension, and may not be used to encode any instruction operand. The content of the register index field 744 directly specifies or specifies through address generation the location of the source or destination operand (in a register or in memory). In some embodiments, the register field 744 is supplemented with additional bits from a prefix (e.g., prefix 601) to allow for greater addressing.
[0077] The R / M field 746 can be used to encode an instruction operand that references a memory address, or can be used to encode a destination register operand or a source register operand. Note that the R / M field 746 can be combined with the MOD field 742 in some embodiments to specify an addressing mode.
[0078] The SIB byte 704 includes a scale field 752, an index field 754, and a base field 756 for address generation. The scale field 752 indicates a scaling factor. The index field 754 specifies the index register to be used. In some embodiments, the index field 754 is supplemented with additional bits from a prefix (e.g., prefix 601) to allow for greater addressing. The base field 756 specifies the base register to be used. In some embodiments, the base field 756 is supplemented with additional bits from a prefix (e.g., prefix 601) to allow for greater addressing. In practice, the content of the scale field 752 allows the content of the index field 754 to be scaled for memory address generation (e.g., for address generation using 2 缩放 * index + base).
[0079] Some addressing forms utilize displacement values to generate memory addresses. For example, it can be based on 2 缩放*Index + base + displacement, index * scale + displacement, r / m + displacement, instruction pointer (RIP / EIP) + displacement, register + displacement, etc. to generate a memory address. The displacement can be a value of 1 byte, 2 bytes, 4 bytes, etc. In some embodiments, the displacement field 607 provides this value. Additionally, in some embodiments, the displacement factor usage is encoded in the MOD field of the addressing field 605, which indicates a compressed displacement scheme for which the displacement value is calculated by multiplying disp8 by a scaling factor N, which is determined based on the vector length, the value of b bits, and the input element size of the instruction. The displacement value is stored in the displacement field 607.
[0080] In some embodiments, the immediate field 609 specifies an immediate value for the instruction. The immediate value can be encoded as a 1 - byte value, 2 - byte value, 4 - byte value, etc.
[0081] Figure 8 An embodiment of the first prefix 601(A) is illustrated. In some embodiments, the first prefix 601(A) is an embodiment of the REX prefix. Instructions using this prefix can specify general - purpose registers, 64 - bit packed data registers (e.g., single instruction, multiple data (SIMD) registers or vector registers), and / or control registers and debug registers (e.g., CR8 - CR15 and DR8 - DR15).
[0082] Instructions using the first prefix 601(A) can specify up to three registers using a 3 - bit field, depending on the format: 1) using the reg field 744 and the R / M field 746 of the Mod R / M byte 702; 2) using the Mod R / M byte 702 and the SIB byte 704, including using the reg field 744, the base field 756, and the index field 754; or 3) using the register field of the opcode.
[0083] In the first prefix 601(A), bit positions 7:4 are set to 0100. Bit position 3 (W) can be used to determine the operand size, but may not determine the operand width alone. Thus, when W = 0, the operand size is determined by the code segment descriptor (CS.D), and when W = 1, the operand size is 64 bits.
[0084] Note that adding another bit would allow addressing 16 (2 4 ) registers, while the individual MOD R / M reg field 744 and MOD R / M R / M field 746 can each address only 8 registers.
[0085] In the first prefix 601(A), bit position 2 (R) can be an extension of the MOD R / M reg field 744 and can be used to modify the ModR / M reg field 744 when this field encodes a general-purpose register, a 64-bit packed data register (e.g., an SSE register), or a control or debug register. When the Mod R / M byte 702 specifies some other register or defines an extended opcode, R is ignored.
[0086] Bit position 1 (X) The X bit can modify the SIB byte index field 754.
[0087] Bit position B (B) B can modify the base address in the Mod R / M R / M field 746 or the SIB byte base address field 756; or it can modify the opcode register field for accessing a general-purpose register (e.g., general-purpose register 525).
[0088] Figures 9A through 9D An embodiment is illustrated of how the R, X, and B fields of the first prefix 601(A) are used. Figure 9A An illustration is shown of how, when the SIB byte 704 is not used for memory addressing, R and B from the first prefix 601(A) are used to extend the reg field 744 and the R / M field 746 of the MOD R / M byte 702. Figure 9B An illustration is shown of how, when the SIB byte 704 is not used, R and B from the first prefix 601(A) are used to extend the reg field 744 and the R / M field 746 of the MOD R / M byte 702 (register-register addressing). Figure 9B An illustration is shown of how, when the SIB byte 704 is used for memory addressing, R, X, and B from the first prefix 601(A) are used to extend the reg field 744, the index field 754, and the base address field 756 of the MOD R / M byte 702. Figure 9D An illustration is shown of how, when a register is encoded in the opcode 603, B from the first prefix 601(A) is used to extend the reg field 744 of the MOD R / M byte 702.
[0089] Figures 10A through 10BIllustrates an embodiment of a second prefix 601(B). In some embodiments, the second prefix 601(B) is an embodiment of the VEX prefix. The second prefix 601(B) encoding allows an instruction to have more than two operands and allows SIMD vector registers (e.g., vector / SIMD register 2410) to be longer than 64 bits (e.g., 128 bits and 256 bits). The use of the second prefix 601(B) provides a syntax for three-operand (or more) operations. For example, a previous two-operand instruction performed an operation such as A = A + B, which overwrote the source operand. The use of the second prefix 601(B) enables the operands to perform non-destructive operations, such as A = B + C.
[0090] In some embodiments, the second prefix 601(B) has two forms - a two-byte form and a three-byte form. The two-byte second prefix 601(B) is mainly used for 128-bit, scalar, and some 256-bit instructions; while the three-byte second prefix 601(B) provides a compact replacement for the first prefix 601(A) and 3-byte opcode instructions.
[0091] Figure 10A Illustrates an embodiment of the two-byte form of the second prefix 601(B). In one example, the format field 1001 (byte 0 1003) contains the value C5H. In one example, byte 1 1005 includes an "R" value in bit [7]. This value is the complement of the same value of the first prefix 601(A). Bit [2] is used to specify the length (L) of the vector (where a value of 0 is scalar or a 128-bit vector, and a value of 1 is a 256-bit vector). Bits [1:0] provide opcode extensibility equivalent to some traditional prefixes (e.g., 00 = no prefix, 01 = 66H, 10 = F3H, and 11 = F2H). Bits [6:3], shown as vvvv, can be used for: 1) encoding the first source register operand, which is specified in inverted (one's complement) form and is valid for instructions with two or more source operands; 2) encoding the destination register operand, which is specified in one's complement form and is used for certain vector shifts; or 3) not encoding any operand, and this field is reserved and should contain a certain value, such as 1111b.
[0092] Instructions using this prefix can use the Mod R / M R / M field 746 to encode instruction operands that reference a memory address, or encode a destination register operand or a source register operand.
[0093] Instructions using this prefix can use the Mod R / M reg field 744 to encode a destination register operand or a source register operand, which is treated as an opcode extension and is not used to encode any instruction operand.
[0094] For instruction syntax that supports four operands, vvvv, the Mod R / M R / M field 746 and the Mod R / M reg field 744 encode three of the four operands. Then bits [7:4] of the immediate 609 are used to encode the third source register operand.
[0095] Figure 10B Illustrated is an embodiment of the three-byte form of the second prefix 601(B). In one example, the format field 1011 (byte 0 1013) contains the value C4H. Byte 1 1015 includes "R", "X", and "B" in bits [7:5], which are the complements of the same values of the first prefix 601(A). Bits [4:0] of byte 1 1015 (shown as mmmmm) include content to encode one or more implicit leading opcode bytes as needed. For example, 00001 means a 0FH leading opcode, 00010 means a 0F38H leading opcode, 00011 means a leading 0F3A opcode, and so on.
[0096] Bit [7] of byte 2 1017 is used similarly to the W of the first prefix 601(A), including helping to determine the operand size that can be promoted. Bit [2] is used to specify the length (L) of the vector (where a value of 0 is a scalar or a 128-bit vector, and a value of 1 is a 256-bit vector). Bits [1:0] provide opcode extensibility equivalent to some traditional prefixes (e.g., 00 = no prefix, 01 = 66H, 10 = F3H, and 11 = F2H). Bits [6:3], shown as vvvv, can be used for: 1) encoding the first source register operand, which is specified in inverted (ones' complement) form and is valid for instructions with two or more source operands; 2) encoding the destination register operand, which is specified in ones' complement form, for certain vector shifts; or 3) not encoding any operand, and this field is reserved and should contain a certain value, such as 1111b.
[0097] Instructions using this prefix can use the Mod R / M R / M field 746 to encode instruction operands that reference a memory address, or to encode a destination register operand or a source register operand.
[0098] Instructions using this prefix can use the Mod R / M reg field 744 to encode a destination register operand or a source register operand, which is treated as an opcode extension and is not used to encode any instruction operand.
[0099] For instruction syntax that supports four operands, vvvv, the Mod R / M R / M field 746 and the Mod R / M reg field 744 encode three of the four operands. Then bits [7:4] of the immediate 609 are used to encode the third source register operand.
[0100] Figure 11 Illustrates an embodiment of the third prefix 601(C). In some embodiments, the first prefix 601(A) is an embodiment of the EVEX prefix. The third prefix 601(C) is a four-byte prefix.
[0101] The third prefix 601(C) can encode 32 vector registers (e.g., 128-bit, 256-bit, and 512-bit registers) in 64-bit mode. In some embodiments, write masks / operation masks (see discussion of registers in previous figures, e.g., Figure 5 ) or predicated instructions utilize this prefix. The operation mask register allows for conditional processing or select control. Operation mask instructions - whose source / destination operands are operation mask registers and treat the contents of the operation mask register as a single value - are encoded using the second prefix 601(B).
[0102] The third prefix 601(C) can encode instruction-class-specific functionality (e.g., packed instructions with "load + operation" semantics can support an embedded broadcast function, floating-point instructions with rounding semantics can support a static rounding function, floating-point instructions with non-rounding arithmetic semantics can support a "suppress all exceptions" function, etc.).
[0103] The first byte of the third prefix 601(C) is the format field 1111, which has a value of 62H in one example. The subsequent bytes are referred to as payload bytes 1115 - 1119 and together form a 24-bit value for P[23:0], providing specific capabilities in the form of one or more fields (detailed herein).
[0104] In some embodiments, P[1:0] of payload byte 1119 is the same as the two lowermost mmmm bits. In some embodiments, P[3:2] is reserved. Bit P[4] (R’) allows access to the upper 16 vector register set when combined with P[7] and the ModR / M reg field 744. When SIB type addressing is not needed, P[6] can also provide access to the upper 16 vector registers. P[7:5] consists of R, X, and B, which are vector register, general register, memory addressing operand specifier modifier bits, and when combined with the ModR / M register field 744 and the ModR / M R / M field 746, allow access to the next set of 8 registers beyond the lower 8 registers. P[9:8] provides opcode extensibility equivalent to some traditional prefixes (e.g., 00 = no prefix, 01 = 66H, 10 = F3H, and 11 = F2H). P
[10] is a fixed value 1 in some embodiments. P[14:11], shown as vvvv, can be used for: 1) encoding the first source register operand, which is specified in inverted (one's complement) form and is valid for instructions with two or more source operands; 2) encoding the destination register operand, which is specified in one's complement form for certain vector shifts; or 3) not encoding any operand, and this field is reserved and should contain a certain value, e.g., 1111b.
[0105] P
[15] is similar to the W of the first prefix 601(A) and the second prefix 611(B), and can be used as an opcode extension bit or an operand size promotion.
[0106] P[18:16] specifies the index of a register in an operation mask (write mask) register (e.g., write mask / predicate register 515). In one embodiment of the present invention, a particular value aaa = 000 has a special behavior, implying that no operation mask is used for a particular instruction (which can be implemented in various ways, including using a hardwired all-ones operation mask or hardware that bypasses the masking hardware). When merged, the vector mask allows any set of elements in the destination to be protected from update during the execution of any operation (specified by the base and enhanced operations); in another embodiment, each element of the destination with the corresponding mask bit having 0 retains its old value. In contrast, the zeroing vector mask allows any set of elements in the destination to be zeroed during the execution of any operation (specified by the base and enhanced operations); in one embodiment, the elements of the destination are set to 0 when the corresponding mask bit has a value of 0. A subset of this functionality is the ability to control the vector length of the operation being performed (i.e., the span of the elements being modified, from the first to the last); however, the elements being modified do not have to be contiguous. Thus, the operation mask field allows for partial vector operations, including loads, stores, arithmetic, logic, etc. While embodiments of the present invention have been described in which the content of the operation mask field selects which one of several operation mask registers contains the operation mask to be used (thus the content of the operation mask field indirectly identifies the masking to be performed), alternative embodiments allow, as an alternative or in addition, the content of the mask write field to directly specify the masking to be performed.
[0107] P
[19] can be combined with P[14:11] to encode a second source vector register in a non-destructive source syntax that can utilize P
[19] to access the upper 16 vector registers. P
[20] encodes various functions that vary among different classes of instructions and can affect the meaning of the vector length / rounding control specifier field (P[22:21]). P
[23] indicates support for merge-write masking (e.g., when set to 0) or support for zeroing and merge-write masking (e.g., when set to 1).
[0108] An exemplary embodiment of the encoding of registers in instructions using the third prefix 601(C) is detailed in the following table.
[0109]
[0110]
[0111] Table 1: 32-register support in 64-bit mode
[0112] [2:0] Register type General use REG ModR / M reg GPR, vector Destination or source VVVV vvvv GPR, vector Second source or destination RM ModR / M R / M GPR, vector First source or destination Base address ModR / M R / M GPR Memory addressing Index SIB.Index GPR Memory addressing VIDX SIB.Index Vector VSIB memory addressing
[0113] Table 2: Encoded Register Specifiers in 32-Bit Mode
[0114] [2:0] Register type General use REG ModR / M reg k0 - k7 Source VVVV vvvv k0 - k7 Second source RM ModR / M R / M k0-7 First source {k1] aaa <![CDATA[k0 1 -k7]]> Operation mask
[0115] Table 3: Encoding of Operation Mask Register Specifiers
[0116] Program code can be applied to the input instructions to perform the functions described herein and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, a processing system includes any system having a processor, such as a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor.
[0117] The program code can be implemented in a high-level procedural or object-oriented programming language to communicate with the processing system. If desired, the program code can also be implemented in assembly or machine language. In fact, the mechanisms described herein are not limited in scope to any particular programming language. In any case, the language can be a compiled language or an interpreted language.
[0118] Embodiments of the mechanisms disclosed herein can be implemented in hardware, software, firmware, or a combination of such implementations. Embodiments of the present invention can be implemented as a computer program or program code executed on a programmable system including at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.
[0119] One or more aspects of at least one embodiment can be implemented by representative instructions stored on a machine-readable medium that represent various logic within a processor, which when read by the machine, cause the machine to fabricate the logic to perform the techniques described herein. This representation, referred to as an "IP core", can be stored on a tangible machine-readable medium and provided to various customers or manufacturing facilities to be loaded into the fabrication machines that actually fabricate the logic or processor.
[0120] Such a machine-readable storage medium may include, but is not limited to, a non-transitory tangible arrangement of articles manufactured or formed by a machine or device, including storage media such as: hard disks, any other type of disk (including floppy disks, optical disks, compact disk read-only memory (CD-ROM), compact disk rewritable (CD-RW), and magneto-optical disks), semiconductor devices (such as read-only memory (ROM), random access memory (RAM) such as dynamic random access memory (DRAM) and static random access memory (SRAM), erasable programmable read-only memory (EPROM), flash memory, electrically erasable programmable read-only memory (EEPROM), phase change memory (PCM)), magnetic or optical cards, or any other type of medium suitable for storing electronic instructions.
[0121] Accordingly, embodiments of the present invention also include a non-transitory tangible machine-readable medium that contains instructions or contains design data that defines the structural, circuit, device, processor, and / or system features described herein, such as a Hardware Description Language (HDL). Such an embodiment may also be referred to as a program product.
[0122] Emulation (including binary translation, code morphing, etc.)
[0123] In some cases, an instruction converter may be used to convert instructions from a source instruction set to a target instruction set. For example, the instruction converter may translate (e.g., using static binary translation, dynamic binary translation including dynamic compilation), morph, emulate, or otherwise convert the instructions to one or more other instructions to be processed by the core. The instruction converter may be implemented in software, hardware, firmware, or a combination thereof. The instruction converter may be on the processor, off the processor, or part on the processor and part off the processor.
[0124] Figure 12The block diagram compares the use of a software instruction converter to convert binary instructions in a source instruction set into binary instructions in a target instruction set, according to some implementations. In the illustrated embodiment, the instruction converter is a software instruction converter, although alternatively, the instruction converter may be implemented in software, firmware, hardware, or various combinations thereof. Figure 12 Shows that a program in a high-level language 1202 can be compiled using a first ISA compiler 1204 to generate a first ISA binary code 1206, which can be natively executed by a processor 1216 having at least one first instruction set core. A processor 1216 having at least one first ISA instruction set core represents any such processor: such a processor can execute substantially the instruction set of the first ISA instruction set core or a target code version of an application or other software targeted to run on an Intel processor having at least one first ISA instruction set core by compatibly executing or otherwise processing (1) to perform functions substantially the same as those of a processor having at least one first ISA instruction set core in order to achieve substantially the same results as a processor having at least one first ISA instruction set core. The first ISA compiler 1204 represents a compiler operable to generate a first ISA binary code 1206 (e.g., target code), which can be executed on a processor 1216 having at least one first ISA instruction set core with or without additional linking processing. The first ISA compiler 1204 represents a compiler operable to generate a first ISA binary code 1206 (e.g., target code), which can be executed on a processor 1216 having at least one first ISA instruction set core with or without additional linking processing.
[0125] Similarly, Figure 12 Shows that a program in a high-level language 1202 can be compiled using an alternative instruction set compiler 1208 to generate an alternative instruction set binary code 1210, which can be natively executed by a processor 1214 without a first ISA instruction set core. An instruction converter 1212 is used to convert the first ISA binary code 1206 into code that can be natively executed by a processor 1214 without a first ISA instruction set core. This converted code is unlikely to be the same as the alternative instruction set binary code 1210 because an instruction converter capable of doing so is difficult to make; however, the converted code will implement general operations and consist of instructions from an alternative instruction set. Thus, the instruction converter 1212 represents software, firmware, hardware, or a combination thereof, which allows a processor or other electronic device without a first ISA instruction set processor or core to execute the first ISA binary code 1206 through emulation, simulation, or any other process.
[0126] Apparatus and method for controlling and debugging hardware lockstep mismatches
[0127] Hardware lockstep (HWLS) is a reliability and security (RAS) feature in which two processing entities (e.g., cores, microcontrollers, etc.) operate or execute the same operations in a per-cycle modular fashion. In a typical implementation, a comparator tree compares "critical feature" signals from the two entities in each cycle, and any divergence observed in this traffic results in a so-called "mismatch". Subsequently, a fault signal is issued, and a hard machine check error is then logged in the machine check library.
[0128] Current implementations are inefficient in handling these types of fault detections. In most cases, a halt state is triggered via a micro breakpoint, and root cause analysis (RCA) is performed using traditional debugging methods. Any non-deterministic but expected divergence also results in a lockstep difference, leading to frequent shutdowns of the client.
[0129] Embodiments of the present invention utilize a masking mechanism implemented at multiple levels of a hardware comparator structure to address these issues. In these embodiments, the comparator structure is subdivided into multiple groups, each group associated with a specific number of result signals, where each result signal corresponds to a critical feature metric of a processing element / entity (e.g., a processor core, an interconnect, a cache, an interface, or other IP block). The terms "processing entity" and "processing element" are used interchangeably herein to refer to any type of IP block that supports redundant execution such as lockstep execution.
[0130] A combination of group-level mask bits and signal-level mask bits is configured to mask out certain minor or expected faults that would otherwise result in uncorrectable shutdowns. The group-level and signal-level mask bits also provide precise and efficient identification of the source of the fault. In some embodiments, the masking mechanism can be applied to the results produced by the comparator structure before the results are stored for observation.
[0131] These implementations prevent the parallel execution mode (e.g., lockstep mode) from being interrupted due to non-deterministic and expected functional divergences, thus enabling forward progress for the client (e.g., eliminating uncorrectable shutdowns due to minor errors). In some embodiments, the comparator structure is a hardware comparator structure, although the basic principles of the present invention can be implemented in a comparator structure using a combination of hardware and software / firmware.
[0132] These embodiments overcome various deficiencies of current implementations, such as including: (1) the system halts for debugging, a root cause analysis must be performed, and the mismatch issues are fixed before resuming operations (regardless of the nature of the faults encountered); (2) the lack of sufficient data points to perform root cause analysis to identify problems; and (3) the failure to distinguish between fatal and minor differences, such as selectively expected non-deterministic divergences (e.g., variations in temperature, voltage, etc.), which can lead to the interruption of lockstep.
[0133] Masking known mismatch points and providing the ability to collect more mismatch data points over a period of time speeds up the debugging process, is highly scalable, and achieves cost savings by precisely disabling faulty processing entities / components (e.g., by having survivability patches, fuses, etc.). Additionally, with these techniques, there is no longer a need to gate lockstep progress at the system level due to non-deterministic and non-fatal divergences. These techniques can also be utilized in the manufacturing, testing / sorting of silicon pins. For example, these techniques can be used to mask minor or spurious faults (if any) generated during the execution of test programs.
[0134] In one embodiment, a hardware lockstep mode is implemented for a pair of identical processing elements (e.g., IP blocks (cores, microcontrollers, etc.)), where all functional behaviors are expected to be cycle-accurate between the two entities. A synchronization verification point is established at the traffic intersection of the two entities. Although the embodiments are described in the context of a "lockstep" arrangement, the basic principles of the present invention can be implemented in any parallel execution environment where two (or more) processing entities (e.g., IP blocks such as cores) produce independent results for comparison. These embodiments of the present invention can be used on processors with standard (or non-standard) lockstep configurations, various forms of redundant execution arrangements (including software-based or partially software-based redundancy), and split-lock configurations, where in the split-lock configuration, multiple processor entities perform lockstep processing when in lockstep mode and independent processing when in the "split" mode.
[0135] Figure 13An example embodiment is illustrated in which signals 1301A - 1303A and 1301B - 1303B (e.g., execution results or associated information) respectively generated by a first and a second entity are arranged in groups, illustrated as group blocks 1301 - 1303. Comparators 1310 - 1312 respectively associated with group blocks 1301 - 1303 compare each signal 1301A - 1303A of entity 0 with the corresponding signal 1301B - 1303B of entity 1. Specifically, comparator 1310 compares signal 1301A with the corresponding signal 1301B of group block 1301; comparator 1311 compares signal 1302A with the corresponding signal 1302B of group block 1302; and comparator 1312 compares signal 1303A with the corresponding signal 1303B of group block 1303.
[0136] Each comparator 1310 - 1312 generates the results 1320 - 1322 of the respective comparisons, indicating whether the signals being compared are different or equal. In the illustrated example, most of the results 1320 - 1322 indicate that the signals of entity 0 in each block 1301 - 1303 are equal to the corresponding signals of entity 1. However, one particular result 1330 within the set of group 2 results 1321 indicates that the signal of entity 0 is different from the corresponding signal of entity 1, which typically results in a "mismatch" signal.
[0137] Thus, the total set of signals to be compared is divided into groups, and each group 1301 - 1303 respectively includes a specified number of signals 1301A - B, 1302A - B, 1303A - B. The widths of the various signals may be the same within a group, or may be different within a group or between groups, as long as the corresponding signals being compared between two entities in a group block always have the same width.
[0138] In one embodiment, the orthogonal signal masking scheme is implemented using a number of per-group mask bits (M) 1340 and a number of signal-level mask bits (N) 1350. In this embodiment, the number of signal-level mask bits (N) is equal to the number of signal results (e.g., signal results 1-N 1320 for group 1, signal results 1-N 1321 for group 2, and signal results 1-N 1322 for group M) for each group block 1301-1303. Each corresponding signal result (e.g., result 1330) is associated with a corresponding per-signal mask bit 1350 (e.g., signal-level mask bit 1 is associated with signal result 1 in each group, signal-level mask bit 2 is associated with signal result 2 in each group, and so on). Similarly, the number of per-group mask bits 1340 (M) is equal to the number of groups, and each group block 1301-1303 is associated with one of the per-group mask bits 1340 (e.g., group block 1 is associated with the group 1 mask bit of per-group mask bit 1340, group block 2 is associated with the group 2 mask bit, and so on).
[0139] In operation, each interface gates the corresponding signal-level mask bit of the per-signal mask bits 1350 with the corresponding group-level mask bit of the per-group mask bits 1340 on the path where it creates a mismatch. In one embodiment, this is done at the final stage of the comparator stage to ensure minimal hardware is required for gating. The two sets of mask bits 1340, 1350 create a two-dimensional mask matrix as follows:
[0140]
[0141] Each row of the above matrix corresponds to a different signal mask bit (e.g., the first row corresponds to S1, the second row corresponds to S2, and so on), and each column corresponds to a different group mask bit (e.g., the first column corresponds to G1, the second column corresponds to G2, and so on). The individual bits of the mask matrix are formed by the "AND" of each signal mask bit and the corresponding group mask bit. Thus, the matrix bit is set to 1 only when both the signal mask bit and the corresponding group mask bit are set to 1.
[0142] In one embodiment, the two-dimensional mask matrix is applied to the grouped signal results. When gating is done with both the corresponding signal-level mask bit and the corresponding group mask bit, a mismatch result (FAIL), such as shown for signal result 2 1330 of group 2, can be masked and thus changed to PASS. For example, the masking can be for a specific group-signal combination to filter out minor or otherwise insignificant differences between the results generated by two (or more) processor entities (e.g., cores, microcontrollers, etc.).
[0143] Figure 14 Illustrated is an example result 1400 that includes a set of per-group result bits. The "MISCOMPARE GRP 2" bit 1401 is set to indicate a mismatch if the corresponding mask bit is not set. If the corresponding mask bit is set (e.g., mask bit 2 and group block 2), the fault is considered PASS, as Figure 14 shown. The exact interface that is fired (signal result 2, group 2 in this example) can be identified and can be used for root cause analysis. For example, root cause analysis can include sequentially setting signal-level mask bits until a FAIL is observed.
[0144] The minimum amount of hardware can be used for the described embodiments. Considering Figure 13 as a reference, assuming the number of groups M = 10 and the number of signals per group N = 5, the details of the hardware savings are as follows. Without the described comparator structure, the number of mask bits required would be 10 groups × 5 signals per group, i.e., 50 mask bits. This saves the number of iterations for performing root cause analysis of the interface, but this will become a bottleneck for hardware and scalability when more interfaces are added in the future.
[0145] In contrast, with the comparator structure of embodiments of the present invention, the signal mask bits are shared and are applied orthogonally to all groups. Thus, the number of signal mask bits required is 5 (since there are 5 signals per group), and the number of group mask bits is 10 (unique per group), which means the total number of mask bits is 15, which is easily scalable when new interfaces are introduced in subsequent implementations. The cost to be paid in terms of iterations will be minimal and will be equal to the number of signals plus 1 (i.e., 6 in this example).
[0146] In Figure 15 illustrated is a method according to an embodiment of the present invention. The method can be implemented on various processors and system architectures described herein, but is not limited to any particular architecture.
[0147] At 1501, multiple processor entities (e.g., cores, microcontrollers, etc.) execute program code in parallel to independently generate multiple groups of signals (e.g., results). As previously described, the parallel execution can be performed by the processor entities in a redundant operation mode, such as a static or configurable lockstep mode or other redundant execution modes.
[0148] At 1502, signal comparison is performed within each group of independent signals. Specifically, each signal / result generated by a first processor entity is compared with the corresponding signal / result of a second processor entity. If the comparison indicates a difference in the signal / result, a fault is generated at 1503.
[0149] In one embodiment, at 1504, a matrix mask including a combination of per-signal bit masks and per-group bit masks is applied to a fault. If it is determined at 1505 that the matrix mask indicates that the fault is masked (i.e., based on associated signal and group blocks), then the fault is ignored at 1507 and execution continues in parallel at 1501.
[0150] If the fault is not masked (i.e., the corresponding mask matrix value is not set), then at 1506, the signal-level mask bits in the fault group are used to perform a root cause analysis of the fault. Both the group mask bits and the signal-level mask bits are to be set to accurately pinpoint the exact signal that does not match. For example, the root cause analysis can include sequentially setting the signal-level mask bits until FAIL is observed.
[0151] Embodiments of the present invention provide a scalable and instantiable solution that can be generally applied to any number of compared interfaces, regardless of signal width, and uses minimal hardware. These embodiments help to accurately identify and mask faulty interface signals, thus advancing the process that would otherwise be gated until the first fault is resolved. Using these embodiments, non-deterministic disagreements in the signal / result expectations do not cause disruption of the system-level lockstep mode, resulting in more efficient operation and shorter down times.
[0152] The mask bits can be configured and stored in various ways. For example, the mask bits can be specified in software or firmware and loaded during system startup. Alternatively, or additionally, the mask bits can be configured as fuses or patches on the processor, providing survivability options at the product level and achieving cost savings by reusing existing hardware.
[0153] In some embodiments, debugging of the lockstep mode is performed by dynamically configuring the mask bits in real time (e.g., at runtime) based on observed faults. The faulty interface that caused the mismatch can be marked (e.g., within the MSR or other storage device).
[0154] Figure 16Illustrated is an example where the first core 1601 and the second core 1602 operate in a lockstep mode and generate corresponding signals / results 1620 - 1621. As described above, one embodiment of the comparison circuit / logic 1630 performs a comparison of signals within a signal group and determines whether the result 1650 should reflect any differences in the signals 1620 - 1621 based on the mask matrix 1610. For example, the comparison circuit / logic 1620 may detect a fault condition (i.e., a difference in signals within a specific group), but may not indicate the fault in the result 1650 if the corresponding mask bit from the mask matrix 1610 is set (i.e., the bit associated with the corresponding group and signal). If the corresponding mask bit is not set, the result 1650 will indicate the fault condition.
[0155] Example
[0156] The following are example implementations of different embodiments of the present invention.
[0157] Example 1. A processor, comprising: a plurality of processing elements operable in a redundant processing mode, each of the plurality of processing elements executing the same plurality of instructions and generating corresponding plurality of result signals; a comparator circuit to compare corresponding result signals among the corresponding plurality of result signals, the comparator circuit generating one or more fault indications when a first one or more result signals generated by a first processing element are different from corresponding second one or more result signals generated by a second processing element; and a masking circuit to mask a first fault indication among the one or more fault indications when a corresponding mask bit of a mask matrix is set to a first value.
[0158] Example 2. The processor according to Example 1, wherein the mask matrix is generated based on a combination of signal - level mask bits and group - level mask bits.
[0159] Example 3. The processor according to Example 1 or 2, wherein the comparator circuit includes a plurality of comparators, each of the plurality of comparators being configured to perform a comparison of a group of result signals among the corresponding plurality of result signals, each group of result signals being associated with a group - level mask bit among a plurality of group - level mask bits.
[0160] Example 4. The processor according to any one of Examples 1 - 3, wherein each group of result signals includes result signals of a specific type.
[0161] Example 5. The processor according to any one of Examples 1 - 4, wherein each group of result signals includes result signals associated with individual key characteristic metrics of the plurality of processing elements.
[0162] Example 6. The processor according to any one of Examples 1-5, wherein the corresponding mask bit of the mask matrix is set to the first value only when the corresponding signal-level mask bit among the plurality of signal-level mask bits and the corresponding group-level mask bit among the plurality of group-level mask bits are both set to the first value.
[0163] Example 7. The processor according to any one of Examples 1-6, wherein the first value includes the binary value 1.
[0164] Example 8. The processor according to any one of Examples 1-7, wherein the masking circuit masks the first fault indication by changing the first fault indication to a pass indication.
[0165] Example 9. The processor according to any one of Examples 1-7, wherein the plurality of group-level mask bits and signal-level mask bits are used to identify the source of the one or more fault indications.
[0166] Example 10. A method, comprising: operating a plurality of processing elements in a redundant processing mode, each processing element among the plurality of processing elements executing the same plurality of instructions and generating a corresponding plurality of result signals; comparing the corresponding result signals among the corresponding plurality of result signals, generating one or more fault indications when the first one or more result signals generated by a first processing element are different from the corresponding second one or more result signals generated by a second processing element; and masking a first fault indication among the one or more fault indications when the corresponding mask bit of a mask matrix is set to a first value.
[0167] Example 11. The method according to Example 10, further comprising: generating the mask matrix based on a combination of signal-level mask bits and group-level mask bits.
[0168] Example 12. The method according to Example 10 or 11, wherein the comparison is performed using a plurality of comparators, each comparator among the plurality of comparators being configured to perform a comparison of a group of result signals among the corresponding plurality of result signals, each group of result signals being associated with a group-level mask bit among the plurality of group-level mask bits.
[0169] Example 13. The method according to any one of Examples 10-12, wherein each group of result signals includes result signals of a specific type.
[0170] Example 14. The method according to any one of Examples 10-13, wherein each group of result signals includes result signals associated with individual key characteristic indicators of the plurality of processing elements.
[0171] Example 15. The method according to any one of Examples 10-14, wherein the corresponding mask bit of the mask matrix is set to the first value only when the corresponding signal-level mask bit among the plurality of signal-level mask bits and the corresponding group-level mask bit among the plurality of group-level mask bits are both set to the first value.
[0172] Example 16. The method according to any one of Examples 10-15, wherein the first value includes the binary value 1.
[0173] Example 17. The method according to any one of Examples 10-16, wherein masking the first fault indication includes changing the first fault indication to a pass indication.
[0174] Example 18. The method according to any one of Examples 10-18, further comprising: identifying the source of the one or more fault indications based on the plurality of group-level mask bits and signal-level mask bits.
[0175] Example 19. A machine-readable medium having program code stored thereon, which when executed by a machine, causes the machine to perform operations, the operations including: operating a plurality of processing elements in a redundant processing mode, each of the plurality of processing elements executing the same plurality of instructions and generating corresponding plurality of result signals; comparing the corresponding result signals among the corresponding plurality of result signals, generating one or more fault indications when the first one or more result signals generated by a first processing element are different from the corresponding second one or more result signals generated by a second processing element; and masking a first fault indication among the one or more fault indications when a corresponding mask bit of a mask matrix is set to a first value.
[0176] Example 20. The machine-readable medium according to Example 19, further comprising program code to cause the machine to perform the following operations: generating the mask matrix based on a combination of signal-level mask bits and group-level mask bits.
[0177] Example 21. The machine-readable medium according to Example 19 or 20, wherein the comparison is performed using a plurality of comparators, each of the plurality of comparators being configured to perform a comparison of a group of result signals among the corresponding plurality of result signals, each group of result signals being associated with a group of group-level mask bits among the plurality of group-level mask bits.
[0178] Example 22. The machine-readable medium according to any one of Examples 19-21, wherein each group of result signals includes result signals of a specific type.
[0179] Example 23. The machine-readable medium as described in any one of Examples 19-22, wherein each result signal group includes result signals associated with individual key characteristic metrics of the plurality of processing elements.
[0180] Example 24. The machine-readable medium as described in any one of Examples 19-23, wherein the corresponding mask bit of the mask matrix is set to the first value only when the corresponding signal-level mask bit among the plurality of signal-level mask bits and the corresponding group-level mask bit among the plurality of group-level mask bits are both set to the first value.
[0181] Example 25. The machine-readable medium as described in any one of Examples 19-24, wherein the first value includes the binary value 1.
[0182] Example 26. The machine-readable medium as described in any one of Examples 19-25, wherein masking the first fault indication includes changing the first fault indication to a pass indication.
[0183] Example 27. The machine-readable medium as described in any one of Examples 19-26, further comprising program code to cause the machine to perform the following operations: identify the source of the one or more fault indications based on the plurality of group-level mask bits and signal-level mask bits.
[0184] Embodiments of the present invention may include various steps, which have been described above. These steps may be embodied in machine-executable instructions that can be used to cause a general-purpose or special-purpose processor to execute these steps. Alternatively, these steps may be performed by specific hardware components that include hardwired logic for performing these steps, or by any combination of programmed computer components and custom hardware components.
[0185] As described herein, an instruction can refer to a specific configuration of hardware such as an application specific integrated circuit (ASIC) configured to perform certain operations or having a predetermined function, or software instructions stored in a memory embodied in a non-transitory computer-readable medium. Thus, the techniques illustrated in the figures can be implemented using code and data stored and executed on one or more electronic devices (e.g., a terminal station, a network element, etc.). Such electronic devices utilize computer machine-readable media to store and communicate (internally and / or over a network with other electronic devices) code and data, where the computer machine-readable media are, for example, non-transitory computer machine-readable storage media (e.g., magnetic disks; optical disks; random access memory; read-only memory; flash devices; phase change memory) and transitory computer machine-readable communication media (e.g., electrical, optical, acoustic, or other forms of propagated signals - such as carrier waves, infrared signals, digital signals, etc.). Additionally, such electronic devices typically include a set of one or more processors coupled to one or more other components such as one or more storage devices (non-transitory machine-readable storage media), user input / output devices (e.g., a keyboard, a touch screen, and / or a display), and a network connection. The coupling of the set of processors and the other components is typically through one or more buses and bridges (also referred to as bus controllers). The storage devices and the signals carrying network traffic respectively represent one or more machine-readable storage media and machine-readable communication media. Thus, the storage devices of a given electronic device typically store code and / or data to be executed on the set of one or more processors of the electronic device. Of course, one or more portions of embodiments of the present invention can be implemented using different combinations of software, firmware, and / or hardware. Throughout this detailed description, for illustrative purposes, numerous specific details are set forth in order to provide a thorough understanding of the present invention. However, those skilled in the art will appreciate that the present invention can be practiced without some of these specific details. In some instances, well-known structures and functions are not described in detail so as not to obscure the subject matter of the present invention. Accordingly, the scope and spirit of the present invention should be judged in accordance with the appended claims.
Claims
1. A processor, comprising: a plurality of processing elements operable in a redundant processing mode, each processing element of the plurality of processing elements being operable to execute a same plurality of instructions and generate a corresponding plurality of result signals; a comparator circuit for comparing corresponding result signals of the corresponding plurality of result signals, the comparator circuit for generating one or more fault indications when a first one or more result signals generated by the first processing element are different from a corresponding second one or more result signals generated by the second processing element; as well as A masking circuit is configured to mask a first fault indication of the one or more fault indications when a corresponding mask bit of the mask matrix is set to a first value.
2. The processor of claim 1, wherein: The mask matrix is generated based on a combination of signal-level mask bits and group-level mask bits.
3. The processor according to claim 1 or 2, wherein: The comparator circuit includes a plurality of comparators, each comparator of the plurality of comparators being configured to perform a comparison of a group of result signals of the corresponding plurality of result signals, each group of result signals being associated with a group-level mask bit of a plurality of group-level mask bits.
4. The processor of claim 3, wherein: Each result signal group includes result signals of a specific type.
5. The processor of claim 3, wherein: Each result signal group includes result signals associated with a separate key performance indicator of the plurality of processing elements.
6. The processor of claim 3, wherein: The corresponding mask bit of the mask matrix is set to the first value only when a corresponding signal level mask bit of a plurality of signal level mask bits and a corresponding group level mask bit of the plurality of group level mask bits are both set to the first value.
7. The processor of claim 6, wherein: The first value comprises a binary value of one.
8. The processor according to any one of claims 1 to 7, wherein: The masking circuit masks the first fault indication by changing the first fault indication to a pass indication.
9. A processor according to any one of claims 3 to 8, wherein: The plurality of group-level mask bits and the signal-level mask bits are used to identify sources of the one or more fault indications.
10. A method comprising: operating a plurality of processing elements in a redundant processing mode, each processing element of the plurality of processing elements being operable to execute the same plurality of instructions and generate a corresponding plurality of result signals; comparing corresponding result signals among the corresponding plurality of result signals, generating one or more fault indications when the first one or more result signals generated by the first processing element differ from the corresponding second one or more result signals generated by the second processing element; and A first fault indication of the one or more fault indications is masked when a corresponding mask bit of the mask matrix is set to a first value.
11. The method of claim 10, further comprising: The mask matrix is generated based on a combination of signal level mask bits and group level mask bits.
12. The method according to claim 10 or 11, wherein: The comparison is performed using a plurality of comparators, each of the plurality of comparators being configured to perform a comparison of a group of result signals in the corresponding plurality of result signals, each group of result signals being associated with a group-level mask bit in a plurality of group-level mask bits.
13. The method of claim 12, wherein: Each result signal group includes result signals of a specific type.
14. The method of claim 12, wherein: Each result signal group includes result signals associated with a separate key performance indicator of the plurality of processing elements.
15. The method of claim 12, wherein: The corresponding mask bit of the mask matrix is set to the first value only when a corresponding signal level mask bit of a plurality of signal level mask bits and a corresponding group level mask bit of the plurality of group level mask bits are both set to the first value.
16. The method of claim 15, wherein: The first value comprises a binary value of one.
17. The method according to any one of claims 10 to 16, wherein: Masking the first fault indication includes changing the first fault indication to a pass indication.
18. The method of any one of claims 12 to 17, further comprising: A source of the one or more fault indications is identified based on the plurality of group-level mask bits and the signal-level mask bits.
19. A machine-readable medium having a program code stored thereon, wherein when the program code is executed by a machine, the machine performs an operation, the operation comprising: operating a plurality of processing elements in a redundant processing mode, each processing element of the plurality of processing elements executing the same plurality of instructions and generating a corresponding plurality of result signals; comparing corresponding result signals among the corresponding plurality of result signals, generating one or more fault indications when the first one or more result signals generated by the first processing element differ from the corresponding second one or more result signals generated by the second processing element; and A first fault indication of the one or more fault indications is masked when a corresponding mask bit of the mask matrix is set to a first value.
20. The machine-readable medium of claim 19, further comprising program code to cause the machine to: The mask matrix is generated based on a combination of signal level mask bits and group level mask bits.