Widen memory accesses to aligned addresses for misaligned memory operations

By widening the memory access to the aligned address on the RISC processor and modifying the unaligned address using data processing operations, the problem of abnormalities throwing in the unaligned memory operation in the RISC design is solved, and performance improvement and resource conservation are achieved.

CN113646744BActive Publication Date: 2025-06-17MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080027019.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-04-03
Filing Date
2020-03-25
Publication Date
2025-06-17
Estimated Expiration
2040-03-25

AI Technical Summary

Technical Problem

The RISC design based on the load-storage architecture lacks the general ability to perform specific memory operations at any address, resulting in an exception thrown by an unaligned atomic memory operation.

Method used

The memory access is widened to the aligned address by using the next larger power of two, and the original unaligned address is modified to align with the widened access address using data processing operations of shift, rotation, and bit domain manipulation.

Benefits of technology

Avoid misalignment related failure exceptions, improve processor performance, reduce resource requirements, and support effective operation of non-native applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113646744B_ABST
    Figure CN113646744B_ABST
Patent Text Reader

Abstract

On a processor using a load - store instruction set architecture (ISA) that requires aligned accesses, unaligned atomic memory operations are performed by widening the memory access to an aligned address using the next larger power of two (e.g., a 4 - byte access is widened to 8 bytes, and an 8 - byte access is widened to 16 bytes). Data - processing operations including shifts, rotates, and bit - field manipulations supported by the load - store ISA are utilized to modify only the bytes in the original unaligned address such that the atomic memory operation is aligned with the widened access address. Aligned atomic memory operations using widened accesses avoid fault exceptions associated with unaligned accesses for most 4 - byte and 8 - byte accesses. In the case where a memory access crosses a 16 - byte boundary, exception handling is performed.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Common RISC (Reduced Instruction Set Computer) designs based on load-store architectures (such as ARM, ARM64, and PowerPC) employ an instruction set architecture (ISA) that lacks the general ability to perform specific memory operations at arbitrary addresses. In the case of atomic interlocked operations, for example, these architectures require memory operations to be "naturally aligned", where 4-byte operations are executed at 4-byte aligned addresses, 8-byte operations are executed at 8-byte aligned addresses, and so on. If an atomic access is not naturally aligned, a processor employing a RISC architecture will throw an exception. Summary of the Invention

[0002] On a processor using a load-store instruction set architecture (ISA) that requires aligned access, unaligned atomic memory operations are performed by widening the memory access to an aligned address using the next larger power of two (e.g., a 4-byte access is widened to 8 bytes, and an 8-byte access is widened to 16 bytes). Data processing operations including shift, rotate, and bitfield manipulation supported by the load-store ISA are utilized to modify only the bytes in the original unaligned address so that the atomic memory operation is aligned with the widened access address. The aligned atomic memory operation using widened access avoids the fault exceptions associated with unaligned access for most 4-byte and 8-byte accesses. In the case where a memory access crosses a 16-byte boundary, exception handling is performed.

[0003] In an illustrative example, memory access widening is implemented using an emulator operating on a computing device having a RISC-based processor. The emulator is configured to cooperate with an application executing on the computing device that utilizes unaligned memory operations. For example, the application may initially be written for a CISC (Complex Instruction Set Computer) register memory architecture such as x86, and the processor may use a RISC load-store architecture such as ARM, ARM64, PowerPC, or MIPS.

[0004] The emulator includes a memory widening component and an exception handler. The memory widening component receives unaligned x86 memory operation instructions from an application for primitives that lock shared variables or objects in memory. For example, the instructions can include swapping operand contents from memory / register, swap and add, compare and swap, etc. The memory widening component updates the original access to operate on widened 8-byte or 16-byte aligned memory addresses. The exception handler receives exceptions that are thrown and performs one or more tasks based on the exceptions. For example, the exception handler can handle load-link / store-conditional exceptions by internally emulating instructions to perform memory operations. For example, exceptions are thrown in cases such as when a memory access crosses an 8-byte boundary, a 16-byte boundary, or a memory page boundary (e.g., when load-link / store-conditional is architecturally prohibited, or page table protection blocks access), or when the access is suspect (such as through an instruction using a bad pointer). The exception handler can ensure that the original address rather than the widened aligned address is used in at least some propagated exceptions. In some implementations, this helps ensure that other processes (e.g., logger, debugger, interface, etc.) correctly log and handle the exceptions.

[0005] Advantageously, this memory access widening provides an improvement in the performance of the computing device and reduces the demand for resources that may be scarce on some devices. For example, by avoiding exception handling typically associated with unaligned accesses, the processing cycles on the processor required to perform most 4-byte and 8-byte accesses on widened aligned addresses are reduced. In some implementations, the performance improvement can be on the order of multiple magnitudes (e.g., 5000:1). Typically, an application is configured for multi-threaded execution, whereas in contrast exception handling is serially executed on a single thread and all other application threads are suspended. For example, the load time of an x86 application can also be improved, which can reduce the computational overhead and power requirements for executing an x86 application on ARM hardware. The reduced load time can also improve the efficiency of the human-machine interface and the quality of the user experience on the computing device.

[0006] This memory access widening further improves the operation of the computing device by enhancing the backward compatibility of RISC and load-store architectures, e.g., supporting non-native or legacy applications that utilize unaligned memory access operations. This backward compatibility is useful for both users and platform developers - users can realize the benefits of new or different platforms while continuing to utilize non-native applications, and platform developers are able to meet the needs of a larger user base, which can encourage the development of more diverse and rich platform features.

[0007] The Summary of the Invention introduces a series of concepts in a simplified form, and these concepts will be further described in the detailed description below. The Summary of the Invention is not intended to identify the key features or essential features of the claimed subject matter, nor is it intended to assist in determining the scope of the claimed subject matter. In addition, the claimed subject matter is not limited to implementations that solve any or all of the disadvantages noted in any part of this disclosure. It should be understood that the above subject matter can be implemented as a computer-controlled device, a computer process, a computing system, or as an article of manufacture such as one or more computer-readable storage media. These and various other features will become apparent by reading the following detailed description and referring to the related drawings. Brief Description of the Drawings

[0008] Figure 1 Illustrates an exemplary computing architecture;

[0009] Figure 2 Illustrates exemplary aligned and unaligned memory operations;

[0010] Figure 3 Illustrates an exemplary embodiment of the present memory access widening to an aligned address for an unaligned memory operation;

[0011] Figure 4 Illustrates an exemplary emulator configured to perform the present memory access widening to an aligned address for an unaligned memory operation;

[0012] Figure 5 Illustrates the operational details of an exemplary emulator;

[0013] Figure 6 Is an illustrative flowchart of memory access widening for a 4-byte access;

[0014] Figure 7 Illustrates an exemplary embodiment of memory access widening for a 4-byte access;

[0015] Figure 8 and Figure 9 Illustrates an exemplary embodiment of memory access widening for a 4-byte access that crosses an 8-byte boundary;

[0016] Figure 10 Is an illustrative flowchart of memory access widening for an 8-byte access;

[0017] Figure 11 and 12 Illustrates an exemplary embodiment of memory access widening for an 8-byte access;

[0018] Figure 13 Illustrates the operational details of an exemplary exception handler;

[0019] Figure 14 illustrates an illustrative computing environment in which a remote emulation service may operate;

[0020] Figure 15 and 16 17 are flowcharts of illustrative methods for memory access widening to aligned addresses for unaligned memory operations;

[0021] Figure 18 is a block diagram of an illustrative computing device architecture that can be used, at least in part, to implement the present memory access widening to aligned addresses for unaligned memory operations;

[0022] Figure 19 is a simplified block diagram of an illustrative computing device that can be used, at least in part, to implement the present memory access widening to aligned addresses for unaligned memory operations.

[0023] In the drawings, like reference numerals denote like elements. Unless otherwise noted, the elements are not drawn to scale. DETAILED DESCRIPTION

[0024] In computational science, an atomic access is an attempt to perform an exclusive read, write, or modification of shared data in a storage device on a computing device. By default, most reads and writes are atomic. However, in some ISAs, there are specific commands that can ensure this atomicity. For example, the x86 family of ISAs uses the LOCK prefix on atomic read-modify-write commands. A read-modify-write command is an instruction that combines a read and arithmetic and writes the result. The ARM family of ISAs uses a load-link / store conditional command pair called load / store exclusive (LDXR / STXR) to implement aligned memory access with atomicity. Other architectures may use, for example, compare-and-swap instructions for atomic memory access. These atomic accesses occur in computer programs designed to run on a specific processor. For example, x86 processors are called register-memory architectures, and ARM processors are called load-store architectures. It may be desirable to run a computer program generated for one processor on another processor using a different architecture, for example, to enhance backward compatibility of the hardware and forward compatibility of the application.

[0025] Processors are designed to operate based on their own prescribed ISA, which can program computing devices. Processors can thus be programmed to perform any number of functions based on applications written for a particular ISA. As a result, an application written for one ISA (i.e., configured to run on one computer processor) is difficult to install and run on a processor operating based on a different ISA. For example, as discussed in more detail below, the x86 ISA supports unaligned memory access for atomic operations, while the ARM ISA requires aligned access, and attempts at unaligned access will fail, which are mitigated based on exceptions. When using unaligned memory access for atomic operations, the widening of this memory access to an aligned address for unaligned memory operations advantageously supports non-native applications on a computing device using a load-store ISA to run on the processor with a similar performance as when using native applications.

[0026] Turning now to the drawings, Figure 1 An illustrative computing architecture 100 operating on a computing device 110 employed by a user 105 is shown. The architecture is hierarchical and includes a user mode 115, a kernel mode 120, and a hardware layer 125. The user mode includes applications 130 and a set of application programming interfaces (APIs) 135 with which the applications communicate. The kernel mode is typically instantiated as part of an operating system (OS) kernel operating on the computing device. The kernel includes kernel APIs 140 that arbitrate access to kernel functions, and various kernel functions 145, such as I / O (input / output) 150, security 155, display control 160 (e.g., accessing a display or monitor), memory management 165, and other privileged / kernel functions 170. The kernel is typically configured to have exclusive access to the hardware in the hardware layer 125 through device drivers 172 and other hardware interfaces (I / F) 174.

[0027] The hardware layer 125 is below the kernel mode 120 and includes the hardware of the actual computing device 110. For example, the hardware includes one or more processors 176, memory 178, I / O 180, other hardware 182, etc. In this illustrative example, the processor operates as a load-store processor under a RISC or other suitable ISA that employs aligned access for atomic memory operations, which are invoked by applications 130 or other functions and processes executing on the computing device. In an alternative implementation, the processor can be configured to use compare and swap instructions for atomic operations. The processor is configured with one or more cores 184. Each core typically includes a plurality of arithmetic logic units (ALUs) 186 and associated registers 188 and a controller 190.

[0028] Figure 2Illustrative aligned and unaligned memory operations are shown. In memory 205, the 4-byte access 210 at bytes 4 through 7 is naturally aligned with the corresponding four bytes of the memory. "Natural alignment" means that the access object is aligned with at least one multiple of its own size. Thus, each N-byte access must be aligned on an N-byte memory address boundary (e.g., the address % N (addr % N) must be zero). In contrast, in memory 215, the 4-byte access 220 is unaligned with the memory and straddles an 8-byte boundary, as indicated by reference numeral 225. As described above, the processor 176 ( Figure 1 ) throws an exception when processing an unaligned access.

[0029] This problem is solved by widening the memory access, as Figure 3 illustrated. In the 4-byte access example (as indicated by reference numeral 305), 4 bytes are placed in a larger 8-byte block (as indicated by reference numeral 310). Then the 8-byte block write is widened to the corresponding memory address 315, which has been widened to 8 bytes. Since the 8-byte block is naturally aligned with the widened memory access, the processor will process this operation as if it were an instruction from the native ISA.

[0030] In the 8-byte access example (as indicated by reference numeral 320), 8 bytes are placed in a larger 16-byte block (as indicated by reference numeral 325). Then the 16-byte block is written to the corresponding memory address 330, which has been widened to 16 bytes. This memory access widening solution can also be applied to scenarios where the processor imposes memory size requirements beyond natural alignment. For example, when the processor imposes a 4-byte alignment requirement on an unaligned 2-byte memory access, memory widening can be utilized. In this case, the unaligned 2-byte access can be widened to 4 bytes so that the processor can process the memory operation without throwing an exception. Memory access widening can also be utilized in cases where the access is widened beyond the native register width supported by a given ISA.

[0031] Figure 4 An illustrative emulator 405 is shown, which is configured to implement the widening of memory accesses to aligned addresses for current unaligned memory operations. The emulator is part of a computing device architecture 400, which is similar to Figure 1The architecture shown in and described in the accompanying text. The architecture is layered and includes a user mode 410, a kernel mode number 415, and a hardware layer 420. Native applications 425 interact with native APIs 430 in the user mode, and the native APIs 430 interface with native kernel APIs 435 in the kernel mode. Non-native applications 440 interact with non-native APIs 445 in the user mode. Architecture 400 also supports device drivers 455, hardware I / F 460, processor 465, memory 470, I / O 475, and other hardware 480.

[0032] The emulator 405 can be configured to operate in the user mode 410 or the kernel mode 415, or any combination of the user mode and the kernel mode, as Figure 4 illustrated illustratively. The emulator interfaces between the non-native API 445 and the kernel functions 450 in the kernel mode. As described in more detail below, the emulator includes functions configured for memory access widening and exception handling. In an alternative arrangement, these functions can be implemented in native code in architecture 400 or by using external resources such as a compiler 485. The compiler can be configured to compile non-native application code that is arranged to perform one or more of the emulation functions described in the text. In some other alternative arrangements, some or all of the functions of the emulator can be executed remotely or in some combination of local processing and remote processing.

[0033] Figure 5 Illustrative operational details of the emulator 405 are shown. As shown, the emulator includes a memory access widening component 510 and an exception handler 515. The emulator receives instructions 520 for unaligned memory operations from the non-native application 440 ( Figure 4 ) space (e.g., an application written according to a register-memory architecture, such as the x86 ISA). The instructions illustratively include 4-byte atomic access operations 525, and the 4-byte atomic access operations 525 include, for example, exchange 530, exchange and add 535, compare and exchange 540, and other 4-byte instructions 545. In the x86 ISA, these instructions would include XCHG, XADD, and CMPXCHG, respectively. The instructions 520 further illustratively include 8-byte atomic access operations 550, and the 8-byte atomic access operations 550 include compare and exchange byte 555 and other 8-byte instructions 560. In the x86 ISA, the instruction 555 is CMPXCHG8B.

[0034] The emulator 405 implements compliance with the processor 465 ( Figure 4The load-store architecture aligned memory instruction 565 of (). The instruction implements various data processing operations 570, which illustratively include arithmetic and logic 575, move and shift 580, bit field manipulation 585, load / store operations 590, and other operations 595. The emulator is also configured to receive a memory access exception 505, which may be thrown during a memory access operation. Exception handling is described in more detail in the attached text below Figure 13 The exception handling is described in more detail in the attached text below.

[0035] Figure 6 is an illustrative flowchart 600 of memory access widening for 4-byte accesses, executed by the emulator 405 ( Figure 4 ) using an ISA with 8-byte registers. The memory access widening process begins at block 605. At block 610, an atomic access operation is received from a non-native application 440 ( Figure 4 ) that includes an original address with arbitrary alignment. The emulator checks the memory access alignment at block 615 by checking if the lower 2 bits of the original address are a zero value. At decision block 620, if the access is aligned, the original 4-byte atomic access operation is executed at block 625. In some cases (as described below), an exception may be thrown when this operation is executed. If so, the exception handler 515 is invoked ( Figure 5 ). The process ends at block 630.

[0036] If the access is not aligned at decision block 620, control is passed to decision block 635, where the emulator 405 determines if the access crosses an 8-byte boundary by determining if bit 2 in the original address is set. If the access does not cross an 8-byte boundary, control is passed to block 640, where the original address is rounded down. In this illustrative example, the original address is rounded down to a multiple of 8 to determine the aligned address (e.g., AlignedAddress = OriginalAddress & 0xfffffff8 (aligned address = original address & 0xfffffff8); alternative code could be AlignedAddress = OriginalAddress & (~((size_t)alignment_required - 1) (aligned address = original address & (~((size_t)alignment_required - 1))). More generally, in response to the original address crossing a 2^W boundary in memory, the aligned address can be rounded down to a multiple of 2^W, where W is the byte size of the largest atomic memory operation executable on the processor.

[0037] At box 645, all 8 bytes at the aligned address are loaded into a temporary register. At box 650, the 4 bytes corresponding to the original access are updated in the temporary register (e.g., by the operation of an x86 XCHG instruction). At box 655, using, for example, a store-conditional instruction STXR, all eight bytes of the temporary register are written to the aligned address, and the process ends at box 630. In some implementations, the values at the byte positions in the temporary register for the widened access at the aligned address (e.g., those bytes excluding the 4 bytes of the original access) may be preserved or stored.

[0038] If it is determined at decision box 635 that the access crosses an 8-byte boundary, control is passed to box 660. The emulator 405 determines whether the access crosses a 16-byte boundary by determining whether bit 3 in the original address is set. If the access crosses a 16-byte boundary, this is considered an edge case, and control is passed to box 625, where a 4-byte operation is performed at its original unaligned address, which can be expected to throw an exception. If the access does not cross a 16-byte boundary, control is passed to box 665, where the original access address is rounded down to a multiple of 16 to determine the aligned address (e.g., AlignedAddress = OriginalAddress & 0xfffffff0 (aligned address = original address & 0xfffffff0)).

[0039] At box 670, using, for example, a load-link instruction LDXR, all 16 bytes at the aligned address are loaded into a pair of 8-byte temporary registers temp1 and temp2. In cases where the ISA uses a native register width greater than 8 bytes, other temporary register configurations may be used. For example, if the native register width is 16 bytes, a single register can be utilized instead of a pair of 8-byte registers as in this illustrative example. At box 675, the middle 8 bytes of the temporary registers temp1 and temp2 are extracted and inserted into a third temporary register temp3. At box 680, the 4 bytes corresponding to the original access are updated in the temporary register temp3 (e.g., by operating an x86 XCHG instruction). At box 685, the updated contents of the temporary register temp3 are inserted back into the middle of the temporary registers temp1 and temp2. At box 690, using, for example, a store-conditional instruction STXR, all 16 bytes of the temporary registers temp1 and temp2 are written to the aligned address. The process ends at box 630. In some implementations, the values at the byte positions in the temporary registers temp1 and temp2 for the widened access at the aligned address (e.g., those bytes excluding the 4 bytes of the original access) may be preserved or stored.

[0040] Figure 7 An illustrative example of memory access widening for 4-byte access is shown. In this example, the original access address 705 is 1 byte unaligned (illustrative values are indicated in bold and correspond to byte positions). Rounding down the original address indicates the aligned address starting from byte 0. For example, using the load-link instruction LDXR (as shown, the order of values is flipped in the register), all 8 bytes (values 0 - 7) at the aligned address are loaded into the temporary register 710.

[0041] Shift and rotate operations are performed under the load-store ISA such that the register is changed as indicated by reference numeral 710. Then the values in the original access (e.g., through the operation of the XCHG instruction) are updated, as indicated by reference numeral 715. Another shift and rotate operation is performed on the temporary register to shift the updated values to their corresponding positions in the original access, as indicated by reference numeral 720. Then all 8 bytes of the temporary register are written to the widened aligned address in memory, as indicated by reference numeral 725. For example, using the corresponding store-conditional instruction (e.g., STXR), the write to memory is implemented.

[0042] Figure 8 and Figure 9 An illustrative example of memory access widening for 4-byte access spanning an 8-byte boundary is shown. The ISA in this example has 8-byte registers. A memory access of all 16 bytes including the original unaligned address 805 is loaded into a pair of temporary registers temp1 and temp2, indicated by reference numerals 810 and 815 respectively. In cases where the ISA uses a native register width greater than 8 bytes, other temporary register configurations can be utilized. For example, if the native register width is 16 bytes, a single register can be utilized instead of the pair of 8-byte registers in this illustrative example. Using the bitfield extraction and insertion operations supported by the load-store ISA, the middle 8 bytes shared between the temporary registers temp1 and temp2 (as indicated by reference numeral 820) are extracted and inserted into the temporary register temp3 825.

[0043] As indicated by reference numeral 830, for example using the x86 XCHG instruction, the 4 bytes corresponding to the original access are updated in the temporary register temp3. The content of the temporary register temp3 is inserted back into the middle 8 bytes of the temporary registers temp1 and temp2, as indicated by reference numeral 835. The entire 16-byte content of the combined temporary registers temp1 and temp2 is written to the widened aligned address in memory, as Figure 9 indicated by reference numeral 905 in. For example, using the store-conditional instruction STXR, the write to memory is implemented.

[0044] Figure 10 is an illustrative flowchart 1000 for memory access widening for 8-byte access. The ISA in this example has 8-byte registers. The memory access widening process begins at block 1005. At block 1010, an atomic access operation is received from a non-native application 440( Figure 4 ). The emulator 405( Figure 4 ) checks the memory access alignment at block 1015 by checking whether the lower 3 bits of the original address are a zero value. At decision block 1020, if the access is aligned, then at block 1025, the original 8-byte atomic access operation is performed. In some cases (as described below), an exception may be thrown when this operation is performed. If so, the exception handler 515 is invoked( Figure 5 ). The process ends at block 1030.

[0045] At decision block 1020, if the access is unaligned, then control is passed to decision block 1035, where the emulator 405 determines whether the original address is 4 modulo 16. If not, then control is passed to block 1025, where an exception is thrown and handled by the exception handler 515( Figure 5 ). If the original address is 4 modulo 16, then control is passed to block 1040, where the emulator 405( Figure 4 ) rounds down the original address to a multiple of 16 to determine the aligned address. At block 1045, all 16 bytes at the aligned address are loaded into a pair of 8-byte temporary registers temp1 and temp2. In cases where the ISA uses a native register width greater than 8 bytes, other temporary register configurations can be utilized. For example, if the native register width is 16 bytes, a single register can be utilized instead of the pair of 8-byte registers in this illustrative example. At block 1050, the middle 8 bytes of the temporary registers temp1 and temp2 are extracted and will be inserted into the temporary register temp3.

[0046] At block 1055, the 8 bytes corresponding to the original access are updated in the temporary register temp3 (e.g., by the operation of the x86 CMPXCHG8B instruction). At block 1060, the updated contents of the temporary register temp3 are inserted back into the middle of the temporary registers temp1 and temp2. At block 1065, all 16 bytes of the temporary registers temp1 and temp2 are written to the aligned address. The process ends at block 1030. In some implementations, the values at the byte positions in the temporary registers temp1 and temp2 for the widened access at the aligned address (e.g., those bytes excluding the 8 bytes of the original access) can be preserved or stored.

[0047] Figure 11 and Figure 12 illustrates an illustrative example of memory access widening for 8 - byte access. In the ISA of this example, there are 8 - byte registers. A memory access including all 16 bytes of the original unaligned address 1105 is loaded into a pair of temporary registers temp1 and temp2, indicated by reference numerals 1110 and 1115 respectively. These load operations can be implemented using load - linked instructions. In cases where the ISA uses a native register width greater than 8 bytes, other temporary register configurations can be utilized. For example, if the native register width is 16 bytes, a single register can be utilized instead of the pair of 8 - byte registers in this illustrative example. As indicated by reference numeral 1120, using bit - field extraction and insertion operations supported by the load - store ISA, the middle 8 bytes shared between the temporary registers temp1 and temp2 are extracted and inserted into the temporary register temp3 1125.

[0048] As indicated by reference numeral 1130, for example using the x86 CMPXCHG8B instruction, the 8 bytes corresponding to the original access are updated in the temporary register temp3. The content of the temporary register temp3 is inserted back into the middle 8 bytes of the temporary registers temp1 and temp2, as indicated by reference numeral 1135. The entire 16 - byte content of the combined temporary registers temp1 and temp2 is written to the widened aligned address in memory, as Figure 12 indicated by reference numeral 1205 in. For example, a store - conditional instruction (e.g., STXR) corresponding to the load - linked instruction used to load all 16 bytes into the temporary registers can be used to implement the write to memory.

[0049] Figure 13 illustrates that it can be in emulator 405( Figure 4) Operational details of the illustrative exception handler 515 instantiated therein. During memory widening, the memory access widening component 510 uses the associated load-link / store-conditional instructions (e.g., ldl_l / stl_c or ldq_l / stp_c for Alpha, lwarx / stwcx or ldarx / stdcx for PowerPC, ll / sc for MIPS, ldrex / strex for ARMv6 / v7, ldxr / stxr for ARMv8, lr / sc for RISC-V, etc.) to read and write register contents for the widened memory address 1305. For each write attempt, the memory access widening component stores the original access address 1310 as a thread-local variable 1315. The exception handler receives the thrown exception and performs one or more tasks based on the exception. For example, the exception handler is configured to manage load-link / store-conditional exceptions 1320, which can be thrown during memory widening operations when a non-native application 440 ( Figure 4 ) attempts to access memory on the load-store computing device.

[0050] The exception handler 515 can handle the load-link / store-conditional exception 1320 by internally emulating instructions to perform memory operations from the application, as indicated by reference numeral 1322. Other exceptions 1324 from access and / or protection faults can also be thrown, which propagate beyond the boundaries of memory widening and the emulated instruction operations. For example, such exceptions can be caused when a memory access crosses an 8-byte boundary, a 16-byte boundary, or a memory page boundary (e.g., when load-link / store-conditional is architecturally prohibited or page table protection blocks access), or when the access is suspect (such as through an instruction using a bad pointer). The exception handler can allow such exceptions to go unhandled, as indicated by reference numeral 1323, but can provide a report, as described below.

[0051] For one or two types of exceptions (e.g., memory access and propagation exceptions), the exception handler 515 can retrieve the original access address 1310 stored as a thread-local variable. The exception handler can use the retrieved original access address to ensure that the exception record 1325 reports the original access address instead of reporting an aligned access address. Since faults typically occur from the original access addressing, such reporting can improve the accuracy of both exception handling and other processes that can utilize the exception record. Note that the implementation of the disclosed solution ensures that widened aligned memory accesses do not result in accesses across page table protection boundaries, and thus memory protection faults on widened aligned memory accesses also result in the same type of memory protection fault on the original unaligned memory access (when using the original ISA). Further note that when widened aligned memory accesses cross a 16-byte boundary and due to other fault scenarios including, for example, suspect memory calls, incorrect pointers, etc., the implementation of the disclosed solution can result in a predetermined type of exception being thrown.

[0052] Figure 14 An illustrative computing environment 1400 in which the remote emulation service 1405 is operable is shown. In the computing environment 1400, the same or different users 105 may employ various devices 110 that communicate via a communication network 1415. The devices 110 can, in some cases, support voice telephony capabilities and generally support data consumption applications such as Internet browsing and multimedia (e.g., music, video, etc.) consumption and various other features. The devices 110 can include, for example, user devices, mobile phones, cellular phones, feature phones, tablet computers, and smart phones that users often employ to make and receive voice and / or multimedia (e.g., video) calls, participate in messaging (e.g., text messaging) and email communications, use applications, and access services that consume data, browse the World Wide Web, etc.

[0053] Other types of electronic devices may also be usable within the environment 1400, including handheld computing devices, PDAs (personal digital assistants), portable media players, devices that use headsets and earphones (e.g., Bluetooth-compatible devices), phablet devices (e.g., smart phone / tablet device combinations), wearable computing devices such as head-mounted display (HMD) systems and smart watches, navigation devices such as GPS (Global Positioning System) systems, laptop PCs (personal computers), smart speakers, IoT (Internet of Things) devices, smart appliances, connected car devices, smart home hubs and controllers, desktop computers, multimedia consoles, gaming systems, etc. In the subsequent discussion, the use of the term "device" is intended to cover all devices configured with communication capabilities and capable of connecting to the communication network 1415.

[0054] The various devices 110 in the environment 1400 can support different features, functionality, and capabilities (generally referred to herein as "features"). Some of the features supported on a given device may be similar to the features supported on other devices, while other features may be unique to the given device. The degree of overlap and / or distinction between the features supported on the various devices 110 can vary according to the implementation. For example, some devices 110 can support touch control, gesture recognition, and voice commands, while other devices 110 can support a more limited user interface. Some devices can support video consumption and Internet browsing, while other devices can support more limited media processing and network interface functionality.

[0055] Devices 110 can generally utilize the network 1415 to access and / or implement various user experiences. The network can include any one of a variety of network types and network infrastructures in various combinations or sub - combinations, including cellular networks, satellite networks, IP (Internet Protocol) networks (such as Wi - Fi under IEEE 802.11 and Ethernet networks under IEEE 802.3), the public switched telephone network (PSTN), and / or short - range networks (such as networks). The network infrastructure can be supported by, for example, mobile operators, enterprises, Internet service providers (ISPs), telephone service providers, data service providers, etc.

[0056] The network 1415 can utilize a portion of the Internet (not shown) or include an interface that supports a connection to the Internet, such that the devices 110 can access content provided by various remote or cloud - based application services and websites and present a user experience (not shown). The application services and websites can support a variety of features, services, and user experiences, such as social networking, maps, news and information, entertainment, travel, productivity, finance, etc.

[0057] The remote emulation service 1405 can be configured to provide memory access widening to the computing devices 110 as a remote service. For example, one or more of the computing devices 110 can employ RISC - based processors that do not support unaligned memory access as described above, but non - native application support is still desired. In some cases, for example, a given device may not have sufficient resources to perform memory access widening locally, or such local processing is not desired. In such cases, the features provided by the emulator 405 ( Figure 4 ) can be supported by one or more virtual machines executed as part of the service 1405. In other cases, the emulator 405 can be instantiated locally, or utilize a combination of remote and local processing to implement memory access widening.

[0058] Figure 15FIG. 1500 is a flowchart of an illustrative method for performing unaligned atomic memory operations on shared memory using a processor that takes advantage of aligned accesses to memory. Unless otherwise specified, the methods or steps shown in the flowchart and described in the accompanying text are not limited to a particular order or sequence. Additionally, some of these methods or their steps may occur or be performed simultaneously, and not all methods or steps need to be performed in such an implementation depending on the requirements of a given implementation, and some methods or steps may optionally be utilized.

[0059] At step 1505, an atomic memory operation instruction to be executed is received, the atomic memory operation instruction specifying a raw address for accessing N bytes of memory, where the raw address is arbitrarily aligned. At step 1510, if the raw address is naturally aligned, the atomic memory operation is performed at the raw address. At step 1515, if the raw address is unaligned, the access for the atomic memory operation in the instruction is widened to 2*N bytes. At step 1520, using the widened access, an aligned address for the atomic memory operation is determined. At step 1525, using the widened access, the atomic memory operation is performed at the aligned address.

[0060] Figure 16 FIG. 1600 is a flowchart of an illustrative method that may be performed by a computing device configured to perform unaligned atomic memory operations on shared memory. At step 1605, an instruction is received from an application for accessing memory to perform an atomic memory operation at a raw address. At step 1610, it is checked whether the access at the raw address is naturally aligned or unaligned. At step 1615, if the access is naturally aligned, the atomic memory operation is performed at the raw address. At step 1620, if the access is unaligned, the access is widened to an aligned address. At step 1625, the atomic memory operation is performed at the aligned address using the widened access.

[0061] Figure 17FIG. 1700 is a flowchart of an illustrative method that may be executed by one or more RISC processors in a computing device, the processor executing instructions stored on one or more hardware-based non-transitory computer-readable storage media. At step 1705, an expected atomic memory operation expressed using a CISC instruction including an unaligned address is received. At step 1710, in response to the expected atomic memory operation, the memory access is widened to 8 bytes using a 4-byte access to memory that does not cross an 8-byte boundary. At step 1715, in response to the expected atomic memory operation, the memory access is widened to 16 bytes using a 4-byte access to memory that crosses an 8-byte boundary. At step 1720, for an 8-byte widened memory access, an 8-byte aligned address is created, 8 bytes at the aligned address are loaded into a temporary register, the register is updated according to the CISC instruction, and the updated 8-byte content of the register is written to memory at the 8-byte aligned address. At step 1725, for a 16-byte widened memory access, a 16-byte aligned address is created, 16 bytes at the aligned address are inserted into a pair of 8-byte temporary registers, the middle 8 bytes of the pair of temporary registers are extracted into a third temporary register, 4 bytes corresponding to the expected 4-byte access are updated into the middle 8 bytes according to the CISC instruction, the updated 8-byte content of the third temporary register is inserted back into the pair of temporary registers, and the updated 16-byte content of the pair of temporary registers is written to memory at the 16-byte aligned address.

[0062] Figure 18 FIG. 1800 illustrates an illustrative architecture for a device such as a server that is capable of executing the various components for widening memory access described herein. Figure 18The architecture 1800 shown in [figure] includes one or more processors 1802 (e.g., central processing unit, dedicated AI chip, graphics processing unit, etc.), system memory 1804, including RAM (random access memory) 1806 and ROM (read-only memory) 1808, and a system bus 1810 that operatively and functionally couples the components in the architecture 1800. The basic input / output system that contains basic routines to assist in transferring information between elements within the architecture 1800 (such as during startup) is typically stored in the ROM 1808. The architecture 1800 also includes a mass storage device 1812 that is used to store software code or other computer-executable code utilized to implement applications, file systems, and operating systems. The mass storage device 1812 is connected to the processor 1802 through a mass storage controller (not shown) connected to the bus 1810. The mass storage device 1812 and its associated computer-readable storage medium provide non-volatile storage for the architecture 1800. Although the description of computer-readable storage media included herein relates to mass storage devices such as hard disks or CD-ROM drives, those skilled in the art will understand that the computer-readable storage media can be any available storage media accessible by the architecture 1800.

[0063] By way of example and not limitation, computer-readable storage media can include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules, or other data. For example, computer-readable media includes, but is not limited to, RAM, ROM, EPROM (erasable programmable read-only memory), EEPROM (electrically erasable programmable read-only memory), flash memory or other solid-state storage technologies, CD-ROM, DVD, HD-DVD (high definition DVD), Blu-ray or other optical storage, cassette tapes, magnetic tapes, disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and is accessible by the architecture 1800.

[0064] According to various embodiments, the architecture 1800 can operate in a networked environment using a logical connection to a remote computer through a network. The architecture 1800 can be connected to the network through a network interface unit 1816 connected to the bus 1810. It can be understood that the network interface unit 1816 can also be used to connect to other types of networks and remote computer systems. The architecture 1800 can also include an input / output controller 1818 for receiving and processing inputs from a plurality of other devices, including keyboards, mice, touchpads, touchscreens, control devices such as buttons and switches, or electronic styli (in Figure 18(not shown in the figure). Similarly, the input / output controller 1818 may provide output to a display screen, a user interface, a printer, or other types of output devices (also not shown in Figure 18 the figure).

[0065] It will be appreciated that the software components described herein, when loaded into the processor 1802 and executed, can transform the processor 1802 and the overall architecture 1800 from a general-purpose computing system into a special-purpose computing system customized to support the functionality presented herein. The processor 1802 may be composed of any number of transistors or other discrete circuit elements, which may individually or jointly exhibit any number of states. More specifically, the processor 1802 may operate as a finite state machine in response to executable instructions contained within the software modules disclosed herein. These computer-executable instructions can transform the processor 1802 by specifying how the processor 1802 transitions between states, thereby transforming the transistors or other discrete hardware elements that make up the processor 1802.

[0066] Encoding the software modules presented herein may also transform the physical structure of the computer-readable storage medium presented herein. In different implementations of this description, the specific transformation of the physical structure may depend on various factors. Examples of these factors may include, but are not limited to, the technology used to implement the computer-readable storage medium, whether the computer-readable storage medium is characterized as main storage or auxiliary storage, etc. For example, if the computer-readable storage medium is implemented as a semiconductor-based memory, the software disclosed herein can be encoded on the computer-readable storage medium by transforming the physical state of the semiconductor memory. For example, the software can transform the states of the transistors, capacitors, or other discrete circuit elements that make up the semiconductor memory. The software can also transform the physical state of such components to store data thereon.

[0067] As another example, the computer-readable storage medium disclosed herein may be implemented using magnetic or optical technologies. In such an implementation, when the software presented herein is encoded therein, the software can transform the physical state of the magnetic or optical medium. These transformations may include changing the magnetic properties of a specific location within a given magnetic medium. These transformations may also include changing the physical characteristics or properties of a specific location within a given optical medium to change the optical properties of those locations. Other transformations of the physical medium are possible without departing from the scope and spirit of this description, and the above examples are provided only for the purpose of facilitating discussion.

[0068] As can be understood from the foregoing, many types of physical transformations occur in architecture 1800 to store and execute the software components presented herein. It can also be understood that architecture 1800 may include other types of computing devices, including wearable devices, handheld computers, embedded computer systems, smart phones, PDAs, and other types of computing devices known to those skilled in the art. It is also contemplated that architecture 1800 may not include Figure 18 all of the components shown in Figure 18 , may include Figure 18 other components not explicitly shown in

[0069] Figure 19 or may utilize an architecture that is completely different from the architecture shown inis a simplified block diagram of an illustrative computing device 1900, such as a PC, client machine, or server, by which the present memory access widening to aligned addresses for unaligned memory operations can be implemented. The computing device 1900 includes a processor 1905, a system memory 1911, and a system bus 1914 that couples various system components including the system memory 1911 to the processor 1905. The system bus 1914 can be any of several types of bus structures, including a memory bus or memory controller, a peripheral bus, or a local bus using any one of a variety of bus architectures. The system memory 1911 includes a read only memory (ROM) 1917 and a random access memory (RAM) 1921. A basic input / output system (BIOS) 1925, containing basic routines that help to transfer information between elements within the computing device 1900 during startup, is stored in the ROM 1917. The computing device 1900 may also include a hard disk drive 1928 for reading from and writing to an internally disposed hard disk (not shown), a disk drive 130 for reading from or writing to a removable disk 1933, such as a floppy disk, and an optical disk drive 1938 for reading from or writing to a removable optical disk 1943, such as a CD (compact disk), DVD (digital versatile disk), or other optical medium. The hard disk drive 1928, disk drive 1930, and optical disk drive 1938 are connected to the system bus 1914 via a hard disk drive interface 1946, a disk drive interface 1949, and an optical drive interface 1952, respectively. These drives and their associated computer-readable storage media provide non-volatile storage of computer-readable instructions, data structures, program modules, and other data for the computing device 1900. Although this illustrative example includes a hard disk, a removable disk 1933, and a removable optical disk 1943, in some applications of the present memory access widening to aligned addresses for unaligned memory operations, other types of computer-readable storage media that can store computer-accessible data, such as magnetic tape, flash memory cards, digital video disks, data cartridges, random access memory (RAM), read only memory (ROM), etc., may also be used. Additionally, as used herein, the term computer-readable storage medium includes one or more instances of a medium type (e.g., one or more disks, one or more CDs, etc.). For the purposes of this specification and the claims, the phrase "computer-readable storage medium" and its variants are intended to cover non-transitory embodiments and do not include waves, signals, and / or other transient and / or intangible communication media.

[0070] Multiple program modules can be stored on a hard disk, magnetic disk 1933, optical disk 1943, ROM 1917, or RAM 1921. The program modules include an operating system 1955, one or more application programs 1957, other program modules 1960, and program data 1963. A user can enter commands and information into the computing device 1900 through input devices such as a keyboard 1966 and a pointing device such as a mouse 1968. Other input devices (not shown) can include a microphone, a joystick, a gamepad, a satellite antenna, a scanner, a trackball, a touchpad, a touch screen, a touch-sensitive device, a voice command module or device, a user movement or user gesture capture device, and so on. These and other input devices are generally connected to the processor 1905 through a serial port interface 1971 coupled to the system bus 1914, but can also be connected through other interfaces, such as a parallel port, a game port, or a Universal Serial Bus (USB). A monitor 1973 or other type of display device is also connected to the system bus 1914 via an interface (such as a video adapter 1975). In addition to the monitor 1973, a personal computer typically also includes other peripheral output devices (not shown), such as speakers and printers. Figure 19 The illustrative example shown in FIG. also includes a host adapter 1978, a Small Computer System Interface (SCSI) bus 1983, and an external storage device 1976 connected to the SCSI bus 1983.

[0071] The computing device 1900 is operable in a network environment using a logical connection to one or more remote computers, such as the remote computer 1988. The remote computer 1988 can be selected as another personal computer, a server, a router, a network PC, a peer device, or other common network node, and generally includes many or all of the elements described above with respect to the computing device 1900, although Figure 19 only a single representative remote memory / storage device 1990 is shown in FIG. Figure 19 The logical connections depicted in FIG. include a Local Area Network (LAN) 1993 and a Wide Area Network (WAN) 1995. Such a networking environment is often deployed in, for example, offices, enterprise-wide computer networks, intranets, and the Internet.

[0072] When used in a LAN networking environment, computing device 1900 is connected to local area network 1993 through network interface or adapter 1996. When used in a WAN networking environment, computing device 1900 typically includes a broadband modem 1998, a network gateway, or other components for establishing communications over wide area network 1995 (the Internet). Broadband modem 1998, which can be internal or external, is connected to system bus 1914 via serial port interface 1971. In a networking environment, program modules or portions thereof related to computing device 1900 can be stored in remote memory storage device 1990. Note that Figure 19 the network connections shown are illustrative, and depending on the specific requirements of the application of this memory access widening to aligned addresses for unaligned memory operations, other means of establishing a communication link between computers can be used.

[0073] Various exemplary embodiments of this widening of memory access to aligned addresses for unaligned memory operations are now presented by way of illustration and not as an exhaustive list of all embodiments. Examples include one or more hardware-based non-transitory computer-readable memory devices storing computer-executable instructions that, when executed by one or more RISC (Reduced Instruction Set Computer) processors in a computing device, cause the computing device to: receive an expected atomic memory operation that is expressed using a CISC (Complex Instruction Set Computer) instruction including an unaligned original address; in response to the expected atomic memory operation, widen a 4-byte memory access to 8 bytes using a 4-byte access to memory that does not cross an 8-byte boundary; in response to the expected atomic memory operation, widen a 4-byte memory access to 16 bytes using a 4-byte access to memory that crosses an 8-byte boundary; for an 8-byte widened memory access, create an 8-byte aligned address, load 8 bytes at the aligned address into a temporary register, update the register according to the CISC instruction, and write the updated 8-byte content of the register to memory at the 8-byte aligned address; and for a 16-byte widened memory access, create a 16-byte aligned address, insert 16 bytes at the aligned address into a pair of 8-byte temporary registers, extract the middle 8 bytes of the pair of temporary registers into a third temporary register, update 4 bytes corresponding to the expected 4-byte access into the middle 8 bytes according to the CISC instruction, insert the 8-byte content of the third temporary register back into the pair of temporary registers, and write the 16-byte content of the pair of temporary registers to memory at the 16-byte aligned address.

[0074] In another example, the CISC instructions include one or more of the following: XCHG, XADD, CMPXCHG, or CMPXCHG8B. In another example, the instruction also causes the computing device to: receive an expected atomic memory operation expressed using aligned CISC instructions, the expected atomic memory operation utilizing a 4-byte or 8-byte access to memory, and perform the atomic memory operation at the aligned original address. In another example, the instruction also causes the computing device to: in response to the expected atomic memory operation using an 8-byte access to memory and an original 8-byte access that is 4 modulo 16, create a 16-byte aligned address, insert the 16 bytes at the aligned address into a pair of 8-byte temporary registers, extract the middle 8 bytes of the pair of temporary registers into a third temporary register, update the middle 8 bytes according to the CISC instruction, insert the updated 8-byte content of the third temporary register back into the pair of temporary registers, and write the updated 16-byte content of the pair of temporary registers to memory at the 16-byte aligned address. In another example, the instruction also causes the computing device to: in response to the original address not being 4 modulo 16, perform an 8-byte atomic operation at the original address.

[0075] Further examples include a method for performing an unaligned atomic memory operation on a shared memory using memory that utilizes aligned accesses to memory, the method including: receiving an atomic memory operation instruction to be executed, the atomic memory operation instruction specifying an original address for accessing N bytes of memory, where the original address is arbitrarily aligned; if the original address is naturally aligned, performing the atomic memory operation at the original address; if the original address is unaligned, widening the access for the atomic memory operation in the instruction to 2*N bytes; using the widened access to determine an aligned address for the atomic memory operation; and using the widened access to perform the atomic memory operation at the aligned address.

[0076] In another example, natural alignment includes: N-byte accesses are aligned on an N boundary of the address for the memory. In another example, the received atomic memory operation instructions conform to a register-memory instruction set architecture. In another example, widening is implemented using a data processing operation that conforms to a load-store instruction set architecture, the data processing operation including one of bit field manipulation, shift, or rotate. In another example, the performed atomic memory operation uses load-link / store-conditional instructions to write to the widened memory access. In another example, the method further includes: observing an exception that is thrown but not handled by an exception handler, and using the original address for the reported exception instead of using the aligned address. In another example, in response to the original address crossing a 2^W boundary in the memory, the aligned address is rounded down to a multiple of 2^W, where W is the size in bytes of the largest atomic memory operation executable on the processor. In another example, the method is performed in one of: an application, an operating system, an emulator, a remote service, or a combination thereof.

[0077] Further examples include a computing device configured to perform unaligned atomic memory operations on shared memory, including: at least one processor configured with an instruction set architecture including atomic instructions that require aligned memory access; at least one non-transitory memory; and at least one non-transitory computer-readable storage medium having computer-executable instructions stored thereon that, when executed by the at least one processor, cause the computing device to: receive an instruction from an application for accessing memory to perform an atomic memory operation at an original address; check whether the access at the original address is naturally aligned or unaligned; if the access is naturally aligned, perform the atomic memory operation at the original address; if the access is unaligned, widen the access to an aligned address; and perform the atomic memory operation at the aligned address using the widened access.

[0078] In another example, the application is non-native to the instruction set architecture applicable to the processor. In another example, the application employs a CISC (Complex Instruction Set Computer) instruction set architecture and the processor employs a RISC (Reduced Instruction Set Computer) instruction set architecture, or the processor is configured to utilize compare and swap instructions.

[0079] In another example, the instructions also cause the computing device to implement an exception handler that is configured to perform an atomic memory operation in view of the thrown exception using internal emulation and to log the exception using the original faulting address. In another example, the widened access is the next larger power of two accessed in the original address. In another example, the instructions also cause the computing device to update only the bytes associated with the original address according to the atomic memory operation instruction. In another example, the application is instantiated remotely from the computing device.

[0080] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the above specific features and acts are disclosed as example forms of implementing the claims.

Claims

1. One or more hardware-based non-transitory computer-readable memory devices storing computer-executable instructions that, when executed by one or more RISC (Reduced Instruction Set Computer) processors in a computing device, cause the computing device to: Receive an expected atomic memory operation expressed using a CISC (Complex Instruction Set Computer) instruction including an unaligned original address; In response to the expected atomic memory operation, widen a memory access to 8 bytes using a 4-byte access to memory that does not cross an 8-byte boundary; In response to the expected atomic memory operation utilizing 4-byte accesses to the memory that span an 8-byte boundary, widen the memory access to 16 bytes; For the 8-byte widened memory access, create an 8-byte aligned address, load 8 bytes at the aligned address into a temporary register, update the register according to the CISC instruction, and write the updated 8-byte content of the register to the memory at the 8-byte aligned address; And For the 16-byte widened memory access, create a 16-byte aligned address, insert 16 bytes at the aligned address into a pair of 8-byte temporary registers, extract the middle 8 bytes of the pair of 8-byte temporary registers into a third temporary register, update 4 bytes corresponding to the expected 4-byte access to the middle 8 bytes according to the CISC instruction, insert the 8-byte content of the third temporary register back into the pair of 8-byte temporary registers, and write the 16-byte content of the temporary register pair to the memory at the 16-byte aligned address.

2. The one or more hardware-based non-transitory computer-readable memory devices according to claim 1, wherein the CISC instruction includes one or more of the following: XCHG, XADD, CMPXCHG, or CMPXCHG8B.

3. The one or more hardware-based non-transitory computer-readable memory devices according to claim 1, wherein the instructions further cause the computing device to: receive an expected atomic memory operation expressed using a CISC instruction including an aligned original address using a 4-byte or 8-byte access to the memory, and perform the atomic memory operation at the aligned original address.

4. The one or more hardware-based non-transitory computer-readable memory devices according to claim 4, wherein the instructions further cause the computing device to: in response to the expected atomic memory operation, create a 16-byte aligned address using an 8-byte access to the memory and a raw 8-byte address that is 4 modulo 16, insert the 16 bytes at the aligned address into a pair of 8-byte temporary registers, extract the middle 8 bytes of the pair of 8-byte temporary registers into a third temporary register, update the middle 8 bytes according to the CISC instruction, insert the updated 8-byte content of the third temporary register back into the pair of 8-byte temporary registers, and write the updated 16-byte content of the temporary register pair to the memory at the 16-byte aligned address.

5. The one or more hardware-based non-transitory computer-readable memory devices according to claim 4, wherein the instructions further cause the computing device to: in response to the original address not being 4 modulo 16, perform the 8-byte atomic operation at the original address.

6. A method for emulating atomic memory operations on a processor, the atomic memory operations including read-modify-write commands for shared data in a memory, the processor utilizing aligned accesses to the memory, the method comprising: Receive, at the processor, an atomic memory operation instruction to be executed, the atomic memory operation instruction specifying a raw address for accessing N bytes of the memory, where the raw address is arbitrarily aligned; At the processor, if the raw address is naturally aligned, perform the atomic memory operation including reading, modifying, and writing data at the raw address; At the processor, if the raw address is unaligned, widen the access for the atomic memory operation in the instruction to 2*N bytes; At the processor, use the widened access to determine an aligned address for the atomic memory operation, where the aligned address using the widened access includes the raw address; And At the processor, use the widened access to perform the atomic memory operation at the aligned address, where the atomic memory operation performed by the processor includes: Reading data at the aligned address using the widened access, Modifying the data at the raw address according to the data read at the aligned address using the widened access, and Writing the data back to the aligned address using the widened access, where the modified data is written to the raw address.

7. The method according to claim 6, wherein the natural alignment includes: The N-byte access is aligned on the address boundary of N for the memory.

8. The method according to claim 6, wherein the received atomic memory operation instructions conform to a register-memory instruction set architecture.

9. The method according to claim 6, wherein the widening is implemented using a data processing operation conforming to a load-store instruction set architecture, the data processing operation including one of bit field manipulation, shift, or rotate.

10. The method according to claim 6, wherein the executed atomic memory operation utilizes a load-link / store-conditional instruction to write to the widened memory access.

11. The method according to claim 6, further comprising: Observe an exception that is thrown but not handled by an exception handler, and use the raw address for the exception reported to the log instead of using the aligned address.

12. The method according to claim 6, wherein in response to the original address crossing a 2^W boundary in the memory, the aligned address is rounded down to a multiple of 2^W, where 2^W is the size in bytes of the largest atomic memory operation executable on the processor.

13. The method according to claim 6, wherein the received instruction includes a lock prefix, the lock prefix indicating that the received instruction will be executed atomically by the processor.

14. A computing device configured to emulate unaligned atomic memory operations on shared data in a memory, comprising: A processor, configured with an instruction set architecture including atomic instructions, the atomic instructions requiring aligned memory accesses; At least one non-transitory memory; And At least one non-transitory computer-readable storage medium having computer-executable instructions stored thereon, the computer-executable instructions when executed by the processor cause the processor to: Receive an instruction from an application, the instruction being for accessing the memory to perform an atomic memory operation at an original address, the atomic memory operation including a read-modify-write command for the shared data, Check whether the access at the original address is naturally aligned or unaligned. If the access is naturally aligned, perform the atomic memory operation at the original address, If the access is unaligned, widen the access to an aligned address, wherein the aligned address using the widened access includes the original address, and Perform the atomic memory operation at the aligned address using the widened access, wherein the atomic memory operation performed by the processor includes: Read data at the aligned address using the widened access, modify the data at the original address according to the data read at the aligned address using the widened access, and Write the data back to the aligned address using the widened access, wherein the modified data is written to the original address.

15. The computing device according to claim 14, wherein the application is non-native to the instruction set architecture applicable to the processor.

16. The computing device according to claim 15, wherein the application employs a complex instruction set computer (CISC) instruction set architecture, and the processor employs a reduced instruction set computer (RISC) instruction set architecture or the processor is configured to utilize compare and swap instructions.

17. The computing device according to claim 14, wherein the instruction further causes the computing device to implement an exception handler, the exception handler being configured to perform the atomic memory operation in view of the thrown exception using internal emulation and to log the exception using the original fault-causing address.

18. The computing device according to claim 17, wherein the widened access is the next larger power of two accessed in the original address.

19. The computing device according to claim 18, wherein the instruction further causes the computing device to update only the byte associated with the original address based on the atomic memory operation instruction.

20. The computing device according to claim 14, wherein the received instruction includes a lock prefix, the lock prefix indicating that the received instruction will be executed atomically by the processor.

Citation Information

Patent Citations

  • Instruction and logic used for providing vector loading operation / storage operation using spanning function

    CN106951214A

  • Processors, methods, systems, and instructions to atomically store to memory data wider than a natively supported data width

    CN108701027A