Instruction processing device and instruction processing method

By introducing an address translation unit in the GPU to convert the virtual address of the instruction into a physical address in advance, the problem of excessive execution unit size and global memory bandwidth pressure is solved, and the address translation efficiency and convenience of multi-process switching are improved.

CN114327632BActive Publication Date: 2025-07-18SHANGHAI SENSETIME INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011064561.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-30
Publication Date
2025-07-18
Estimated Expiration
2040-09-30

AI Technical Summary

Technical Problem

When performing data processing tasks, existing GPUs need to convert virtual address to physical address through execution units, resulting in excessive volume of execution units and global memory bandwidth pressure.

Method used

Before the instructions are distributed to the execution unit, the virtual address in the instruction is converted into a physical address through the address translation unit, reducing dependence on the execution unit, avoiding setting of TLB and PTW in the execution unit, and setting a larger cache unit using the address translation unit to improve translation efficiency.

Benefits of technology

It reduces the size of the execution unit, avoids the global memory bandwidth problem caused by multiple concurrent page table queries, improves the efficiency of address translation, and simplifies the multi-process context switching process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114327632B_ABST
    Figure CN114327632B_ABST
Patent Text Reader

Abstract

The present disclosure provides an instruction processing device and an instruction processing method. The instruction processing device includes: an instruction processor, an address translation unit, and a target execution unit. The instruction processor is configured to obtain a first instruction and transmit the first instruction to the address translation unit. The address translation unit is configured to receive the first instruction transmitted by the instruction processor, convert the virtual address carried in the first instruction into a physical address, and obtain a second instruction. The target execution unit is configured to execute the second instruction to obtain an instruction execution result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of virtual-to-physical address translation, and more particularly, to an instruction processing device and an instruction processing method. Background Art

[0002] A Graphics Processing Unit (GPU) is a commonly used device for performing cloud-side Artificial Intelligence (AI) inference and training tasks; due to requirements such as resource virtualization, when a current GPU performs a data processing task, it needs to use the function of virtual-to-physical address translation to convert the virtual address in an instruction into a physical address. Summary of the Invention

[0003] Embodiments of the present disclosure provide at least an instruction processing device and an instruction processing method.

[0004] In a first aspect, an embodiment of the present disclosure provides an instruction processing device, including: an instruction processor, an address translation unit, and a target execution unit;

[0005] The instruction processor is configured to obtain a first instruction and transmit the first instruction to the address translation unit;

[0006] The address translation unit is configured to receive the first instruction transmitted by the instruction processor, convert the virtual address carried in the first instruction into a physical address, and obtain a second instruction;

[0007] The target execution unit is configured to execute the second instruction to obtain an instruction execution result.

[0008] In this way, before the instruction is distributed to a specific execution unit, the address translation unit is used to convert the virtual address carried in the instruction into a physical address; when the instruction is distributed to the execution unit, the virtual address carried in the instruction has been translated into a physical address by the address translation unit, so that the execution unit does not need to perform the address translation process, and thus there is no need to set a TLB and a PTW in the execution unit, reducing the volume of the execution unit; at the same time, there will also be no global memory bandwidth problem caused by multiple execution units performing multi-way concurrent page table queries, reducing the bandwidth pressure on the global memory.

[0009] At the same time, in the embodiments of the present disclosure, the translation of the instruction distributed to the execution unit is implemented by the address translation unit, so that it is not limited by the volume of the execution unit, and a larger cache unit can be set for the address translation unit to improve the efficiency of address translation.

[0010] In addition, in the embodiments of the present disclosure, since multiple instructions read physical addresses through the same PTW, translation requests at the same level can be merged, and the translation of different instructions can be executed out of order, thereby reducing memory access, reducing bandwidth pressure, and improving translation efficiency.

[0011] In addition, since only one cache unit needs to be set for the address translation unit, when switching contexts of multiple processes, only the cache in this cache unit needs to be processed, rather than clearing the caches of all execution units. Therefore, it is more conducive to context switching between multiple users and multiple processes.

[0012] In addition, in the instruction processor, there are multiple instruction queues. Limited by the execution efficiency of the execution units, there will be multiple instructions in the instruction queues. Therefore, when performing address translation through the address translation unit, the latency of physical memory query will be "hidden" by the latency of instruction issuance, reducing the high-latency problem caused by TLB miss (the mapping relationship between virtual addresses and physical addresses does not exist in the cache unit).

[0013] In a possible implementation manner, it further includes: an instruction distribution unit;

[0014] The address translation unit is further configured to transmit the second instruction to the instruction distribution unit;

[0015] The instruction distribution unit is configured to, after receiving the second instruction transmitted by the address translation unit, determine the target execution unit among the multiple execution units for the second instruction, and transmit the second instruction to the target execution unit.

[0016] In this way, the first instruction is converted into a second instruction carrying a physical address through the address unit, and the second instruction is distributed to a specific target execution unit through the instruction distribution unit.

[0017] In a possible implementation manner, when executing the second instruction, the target execution unit is configured to:

[0018] Parse the physical address carried in the second instruction and access the physical address.

[0019] In a possible implementation manner, the address translation unit includes: an instruction parsing subunit, an address conversion subunit, and an instruction conversion subunit;

[0020] Among them, the instruction parsing subunit is connected to the instruction processor and is configured to parse the virtual address from the first instruction;

[0021] The address conversion subunit is configured to determine the physical address corresponding to the virtual address parsed by the instruction parsing subunit;

[0022] The instruction conversion subunit is configured to generate a second instruction based on the physical address determined by the address conversion subunit and the first instruction.

[0023] In a possible implementation manner, the address translation unit includes: a cache unit;

[0024] The address translation unit is configured to query the mapping relationship between the virtual address and the physical address stored in the cache unit, and when there is no mapping relationship corresponding to the virtual address in the cache unit, obtain the physical address corresponding to the virtual address from the memory through page table query.

[0025] In a possible implementation manner, the address translation unit is further configured to store the mapping relationship between the physical address obtained from the memory and the virtual address into the cache unit.

[0026] In a possible implementation manner, when the address translation unit converts the virtual address carried in the first instruction into a physical address to obtain a second instruction, it is configured to:

[0027] Change the target flag bit in the first instruction from a first value to a second value, and replace the virtual address with the physical address to obtain the second instruction;

[0028] Wherein, the first value indicates that the address carried in the instruction is a virtual address; the second value indicates that the address carried in the instruction is a physical address.

[0029] In a second aspect, an embodiment of the present disclosure further provides an instruction processing method, including:

[0030] The instruction processor obtains a first instruction and transmits the first instruction to the address translation unit;

[0031] The address translation unit receives the first instruction transmitted by the instruction processor, converts the virtual address carried in the first instruction into a physical address, and obtains a second instruction;

[0032] The target execution unit executes the second instruction to obtain an instruction execution result.

[0033] In a possible implementation manner, it further includes:

[0034] The address translation unit transmits the second instruction to the instruction distribution unit;

[0035] After receiving the second instruction transmitted by the address translation unit, the instruction distribution unit determines the target execution unit among the multiple execution units for the second instruction, and transmits the second instruction to the target execution unit.

[0036] In a possible implementation manner, the target execution unit executes the second instruction, including: the target execution unit parses the physical address carried in the second instruction and accesses the physical address.

[0037] In a possible implementation manner, the address translation unit receives the first instruction transmitted by the instruction processor, converts the virtual address carried in the first instruction into a physical address, and obtains a second instruction, including:

[0038] The address translation unit parses the virtual address from the first instruction, determines the physical address corresponding to the virtual address, and generates a second instruction based on the physical address and the first instruction.

[0039] In a possible implementation manner, it further includes: the mapping relationship between the virtual address and the physical address stored in the address translation query cache unit, and in the case where there is no mapping relationship corresponding to the virtual address in the cache unit, obtaining the physical address corresponding to the virtual address from the memory through page table query.

[0040] In a possible implementation manner, it further includes:

[0041] The address translation unit stores the mapping relationship between the physical address obtained from the memory and the virtual address into the cache unit.

[0042] In a possible implementation manner, the address translation unit converts the virtual address carried in the first instruction into a physical address, and obtains a second instruction, including:

[0043] Changing the target flag bit in the first instruction from a first value to a second value, and replacing the virtual address with the physical address to obtain the second instruction;

[0044] Wherein, the first value indicates that the address carried in the instruction is a virtual address; the second value indicates that the address carried in the instruction is a physical address.

[0045] To make the above objects, features, and advantages of the present disclosure more obvious and understandable, the following specifically gives preferred embodiments and, in conjunction with the accompanying drawings, makes the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings required for the embodiments will be briefly introduced below. The accompanying drawings herein are incorporated into the specification and form a part of this specification. These accompanying drawings show embodiments consistent with the present disclosure and are used together with the specification to illustrate the technical solutions of the present disclosure. It should be understood that the following accompanying drawings only show some embodiments of the present disclosure and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related accompanying drawings can be obtained based on these accompanying drawings without creative efforts.

[0047] Figure 1 The figure shows a schematic diagram of an instruction processing device provided by an embodiment of the present disclosure;

[0048] Figure 2 The figure shows a schematic diagram of an instruction structure provided by an embodiment of the present disclosure;

[0049] Figure 3 The figure shows a schematic diagram of another instruction processing device provided by an embodiment of the present disclosure;

[0050] Figure 4 The figure shows a schematic diagram of an example of an instruction processing device architecture provided by an embodiment of the present disclosure;

[0051] Figure 5 The figure shows a flowchart of an instruction processing method provided by an embodiment of the present disclosure. Detailed implementation manners

[0052] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only some of the embodiments of the present disclosure, rather than all the embodiments. Usually, the components of the embodiments of the present disclosure described and illustrated herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present disclosure is not intended to limit the scope of the claimed present disclosure, but merely represents selected embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present disclosure without creative efforts fall within the scope of protection of the present disclosure.

[0053] A GPU generally includes: a command processor (CP), and multiple compute units (CUs) connected to the command processor; in the related art, the command processor is used to obtain an instruction stream from a host central processing unit (hereinafter referred to as the host), and determine a target execution unit for each instruction in the instruction stream, and then execute the instruction to the corresponding target execution unit; after receiving the instruction executed by the CP, the CU converts the virtual address carried in the instruction into a physical address. Specifically, when the CU converts the virtual address carried in the instruction into a physical address, it needs to determine the corresponding physical address by querying a translation lookaside buffer (TLB) and / or a page table walk (PTW). In this process, due to the large number of execution units in the GPU; a corresponding TLB and PTW need to be set in each execution unit, resulting in the problem that the volume of each execution unit is too large. At the same time, when the GPU executes a data processing task, it divides the data processing task into multiple subtasks and distributes them to different execution units for execution. When multiple execution units execute the corresponding subtasks, there will be multi-way concurrent page table queries, resulting in a global memory bandwidth problem.

[0054] Based on the above research, the present disclosure provides an instruction processing device that, before the instruction is distributed to a specific execution unit, uses an address translation unit to convert the virtual address carried in the instruction into a physical address; when the instruction is distributed to the execution unit, the virtual address carried in the instruction has been translated into a physical address by the address translation unit, so there is no need for the execution unit to perform the address translation process, and thus there is no need to set a TLB and a PTW in the execution unit, reducing the volume of the execution unit; at the same time, there will be no global memory bandwidth problem caused by multi-way concurrent page table queries of multiple execution units, reducing the global memory bandwidth pressure.

[0055] All the defects existing in the above solutions are the results obtained by the inventors through practice and careful research. Therefore, the process of discovering the above problems and the solutions proposed by the present disclosure for the above problems in the following text should be the contributions made by the inventors to the present disclosure during the process of the present disclosure.

[0056] It should be noted that: similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0057] To facilitate the understanding of this embodiment, first, an instruction processing device disclosed in the embodiments of the present disclosure will be introduced in detail.

[0058] The instruction processing device provided by the embodiments of the present disclosure can be used in a Central Processing Unit (CPU), a GPU, or other instruction processing devices including an instruction processor and an execution unit.

[0059] Taking the application of the instruction processing device provided by the embodiments of the present disclosure to a GPU as an example, the instruction processing device provided by the embodiments of the present disclosure will be described below. However, it can also be applied to other types of instruction processing devices.

[0060] See Figure 1 As shown in the figure, it is a schematic structural diagram of the instruction processing device provided by the embodiments of the present disclosure, including: an instruction processor 10, an address translation unit 20, and a target execution unit 30.

[0061] Among them, the instruction processor 10 is used to obtain a first instruction and transmit the first instruction to the address translation unit 20;

[0062] The address translation unit 20 is used to receive the first instruction transmitted by the instruction processor 10, convert the virtual address carried in the first instruction into a physical address, and obtain a second instruction;

[0063] The target execution unit 30 is used to execute the second instruction to obtain an instruction execution result.

[0064] The instruction processor 10, the address translation unit 20, and the target execution unit 30 will be described in detail below.

[0065] When the instruction processor 10 obtains the first instruction, in the case where the instruction processing device provided by the embodiments of the present disclosure is applied to a GPU, the host converts the source program of the software into machine instructions and stores the machine instructions in the instruction memory of the GPU. The instruction processor 10 reads the instructions from the instruction memory. Here, after the instruction processor 10 reads the instructions from the instruction memory, it decouples the instructions with dependencies, and then transmits the first instruction to the address translation unit 20.

[0066] In the case where the instruction processing device provided by the embodiments of the present disclosure is applied to a CPU, the compiler deployed in the computer is responsible for converting the software program of the software into machine instructions and storing the machine instructions in the instruction memory of the CPU; the instruction processing unit 10 reads the instructions from the instruction memory.

[0067] The address translation unit 20 is responsible for converting the virtual address in the first instruction into a physical address after receiving the first instruction transmitted by the instruction processor 10, and generating a second instruction.

[0068] In a specific implementation,Figure 2 As shown in Figure a, an embodiment of the present disclosure provides a schematic structural diagram of an instruction, including an instruction header, a first flag bit, a first address bit, a second flag bit, a second address bit, and an instruction tail. Among them, the first flag bit is used to indicate whether the address stored in the first address bit is a virtual address or a physical address; the second flag bit is used to indicate whether the address stored in the second address bit is a virtual address or a physical address. Exemplarily, when the flag bit is the first value, it indicates that the address stored in the corresponding address bit is a virtual address; when the flag bit is the second value, it indicates that the address stored in the corresponding address bit is a virtual address.

[0069] As Figure 2 As shown in Figure b, a schematic structural diagram of a first instruction is provided. In this example, the first value is 0 and the second value is 1; the first instruction includes two virtual addresses, and the corresponding first flag bit and second flag bit are both 0, indicating that the addresses included in the first instruction are all virtual addresses; the addresses stored in the first address bit and the second address bit are virtual address 0 and virtual address 1 respectively.

[0070] As Figure 2 As shown in Figure c, a schematic structural diagram of a second instruction is provided, where the first flag bit and the second flag bit are both 1, and the address stored in the first address bit is physical address 0 corresponding to virtual address 0; the address stored in the second address is physical address 1 corresponding to virtual address 1.

[0071] To implement the conversion of the first instruction into the second instruction, refer to Figure 3 As shown, an embodiment of the present disclosure provides a specific structure of an address translation unit 20, including: an instruction parsing subunit 21, an address conversion subunit 22, and an instruction conversion subunit 23.

[0072] Among them, the instruction parsing subunit 21 is connected to the instruction processor 10 and is used to parse the virtual address from the first instruction;

[0073] The address conversion subunit 22 is used to determine the physical address corresponding to the virtual address parsed by the instruction parsing subunit 21;

[0074] The instruction conversion subunit 23 is used to generate a second instruction based on the physical address determined by the address conversion subunit 22 and the first instruction.

[0075] In a specific implementation, when the instruction parsing subunit 21 parses the virtual address from the first instruction, the instruction includes an instruction header, a flag bit, an address bit, and an instruction tail, and the above several types of data are stored in different data bits of the instruction; the length of the instruction and the positions of the virtual address and the flag bit in the instruction can be determined by means of internal table lookup; the instruction parsing unit 21 first determines whether the address stored in the address bit corresponding to the flag bit is a virtual address according to the specific position of the flag bit in the first instruction; in the case of a virtual address, the virtual address is read from the corresponding address bit.

[0076] Here, the instruction parsing subunit 21 includes, for example, a first circuit; the number of signal input terminals included in the first circuit can be equal to the number of data bits included in the first instruction; after the flag bit enters the internal circuit structure, a control signal is output by performing a logical operation on the flag bit, and when the action bit of the control signal is 0, the data in the address bit corresponding to the flag bit is transmitted to the address conversion subunit 22 to complete the process of virtual address parsing and transmitting the virtual address to the address conversion subunit 22.

[0077] When determining the physical address corresponding to the virtual address, the address conversion subunit 22 can, for example, access the cache unit 24 connected to the address conversion subunit, query the mapping relationship between the virtual address and the physical address stored in the cache unit 24, and in the case where there is no mapping relationship corresponding to the virtual address in the cache unit, obtain the physical address corresponding to the virtual address from the memory through page table query and transmit the physical address to the instruction conversion subunit 23.

[0078] Here, since the present disclosure embodiment realizes the translation of the instructions distributed to the execution unit through the address translation unit, it can be not limited by the volume of the execution unit, and a larger cache unit can be set for the address translation unit to improve the efficiency of address translation.

[0079] In addition, since only one cache unit needs to be set for the address translation unit, when performing context switching between multiple processes, it is only necessary to process the cache in this cache unit, rather than clearing the caches of all execution units, so it is more beneficial for context switching between multiple users and multiple processes.

[0080] Here, the address conversion subunit 22 can first query in the cache unit 24 whether there is a target virtual address that is the same as the virtual address parsed from the instruction parsing subunit 21. If it exists, based on the logical relationship between the queried target virtual address and the physical address, the physical address corresponding to the target virtual address is read as the physical address corresponding to the virtual address in the first instruction. If it does not exist, the physical address corresponding to the virtual address is obtained from the memory through page table query PTW.

[0081] Here, since the instruction processor 10 can read multiple instructions simultaneously and transmit multiple instructions to the address translation unit concurrently. Therefore, when the address translation unit translates multiple instructions, since there may be a situation where the physical addresses corresponding to the virtual addresses in multiple instructions are read through the same PTW, the translation requests at the same level can be merged, and the translation of different instructions can be executed out of order, thereby reducing the access to memory, reducing the bandwidth pressure, and improving the translation efficiency.

[0082] In another embodiment of the present disclosure, the address translation unit 20 is further configured to store the mapping relationship between the physical address and the virtual address obtained from the memory into the cache unit.

[0083] Here, the address conversion subunit 22 includes, for example, a second circuit, and the number of signal input terminals of the second circuit is, for example, equal to the number of data bits included in the first instruction. After obtaining the virtual address in the first circuit included in the above instruction parsing subunit 21, the second circuit determines whether there is a target virtual address in the cache unit 24 that is the same as the virtual address in the first instruction through logical operations in the order of storage of each virtual address in the cache unit 24. When there is no target virtual address in the cache unit 24 that is the same as the virtual address in the first instruction, the address conversion subunit 22 uses the bus to obtain the physical address corresponding to the virtual address from the memory 50 through the PTW.

[0084] When the instruction conversion subunit 23 converts the first instruction into the second instruction, for example, it can change the target flag bit in the first instruction from the first value to the second value, and replace the virtual address with the physical address to obtain the second instruction;

[0085] Among them, the first value indicates that the address carried in the instruction is a virtual address; the second value indicates that the address carried in the instruction is a physical address.

[0086] Here, the instruction conversion subunit 23 includes, for example, a third circuit; the third circuit can receive the physical address transmitted by the second circuit in the address conversion subunit 22. In addition, the second circuit also transmits the instruction header, instruction tail, and flag bit in the first instruction to the instruction conversion subunit 23. The third circuit changes the flag bit from the first value to the second value through logical operations, and generates and outputs the second instruction based on the instruction header, the flag bit with the value changed, the physical address, and the instruction tail.

[0087] When the target execution unit 30 executes the second instruction, for example, it can parse the physical address carried in the second instruction and access the physical address.

[0088] For example, at a storage location corresponding to a physical address, an operand corresponding to a second instruction is stored; when the target execution unit 30 accesses the physical address, it can read the operand stored at the storage location corresponding to the physical address; then, using the operand, it executes the specific operation indicated by the second instruction to obtain an instruction execution result.

[0089] As Figure 3 shown, an embodiment of the present disclosure also provides another instruction processing device, further including: an instruction distribution unit 40. The address translation unit 20 is further configured to transmit the second instruction to the instruction distribution unit 40;

[0090] The instruction distribution unit 40 is configured to, after receiving the second instruction transmitted by the address translation unit 20, determine the target execution unit 30 among the multiple execution units for the second instruction, and transmit the second instruction to the target execution unit.

[0091] Referring to Figure 4 shown, an embodiment of the present disclosure also provides a specific example of an instruction processor architecture, including: an instruction processor (Command Processor, CP), an address translation unit, a Peripheral Component Interconnect Express (PCIE) interface, an execution unit 0, an execution unit 1, an instruction distribution unit, a memory, and a bus.

[0092] After the instruction processor 10 obtains an instruction from the host through the bus and PCIE and resolves the dependencies, before sending the instruction to the instruction issuing unit, it transmits the instruction to the instruction parsing subunit in the address translation unit. The instruction parsing subunit parses the instruction to obtain the virtual address contained in the instruction, and transmits the virtual address to the address conversion subunit in the address translation unit.

[0093] The instruction form is as Figure 2 shown, and may contain multiple addresses. It is indicated by a flag bit whether it is a virtual address or a physical real address. Among them, a flag bit of 0 indicates a virtual address, and a flag bit of 1 indicates a physical real address.

[0094] The address conversion subunit queries the TLB; if a corresponding physical address exists in the TLB, it sends the queried physical address to the instruction conversion subunit in the address translation unit. The instruction conversion subunit replaces the virtual address in the instruction with the physical real address to complete the instruction translation process.

[0095] If the corresponding physical address does not exist in the TLB, the address translation subunit needs to query the mapping relationship between the virtual address and the physical address from the page table in the memory through the bus. At this time, during the process of accessing the page table in the memory to obtain the physical address, the instruction processor can still issue new instructions to be parsed and further buffer queried.

[0096] Here, since there are multiple instruction queues in the instruction processor, limited by the execution efficiency of the execution unit, there will be multiple instructions in the instruction queue. Therefore, when performing address translation through the address translation unit, the latency of physical memory query will be "hidden" by the latency of instruction issuance, reducing the high latency problem caused by TLB miss (the mapping relationship between the virtual address and the physical address does not exist in the cache unit).

[0097] After the address translation subunit obtains the physical address by querying the page table, it adds the mapping relationship between the virtual address in the instruction and the queried physical address in the TLB or replaces the old mapping relationship, and sends the queried physical actual address to the instruction conversion subunit.

[0098] The instruction conversion subunit performs the following operations: sets the flag bit of the address with the flag bit of 0 (this is the virtual address) to 1, and replaces the virtual address with the physical address to complete the translation process of the instruction.

[0099] After the instruction conversion subunit completes the instruction translation, it passes the translated instruction to the instruction distribution unit, and the instruction distribution unit distributes the instruction to the execution unit.

[0100] In the embodiment of the present disclosure, before the instruction is distributed to the specific execution unit, the address translation unit is used to convert the virtual address carried in the instruction into a physical address; when the instruction is distributed to the execution unit, the virtual address carried in the instruction has been translated into a physical address by the address translation unit, so there is no need for the execution unit to perform the address translation process, and thus there is no need to set up TLB and PTW in the execution unit, reducing the volume of the execution unit; at the same time, there will also be no global memory bandwidth problem caused by multiple execution units performing multi-way concurrent page table queries, reducing the bandwidth pressure on the global memory.

[0101] Based on the same inventive concept, an instruction processing method corresponding to the instruction processing device is also provided in the embodiment of the present disclosure. Since the principle of solving problems by the device in the embodiment of the present disclosure is similar to that of the above-mentioned instruction processing device in the embodiment of the present disclosure, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.

[0102] Refer to Figure 5 As shown in the flowchart of an instruction processing method provided by the embodiment of the present disclosure, the instruction processing method includes:

[0103] S501: The instruction processor obtains a first instruction and transmits the first instruction to the address translation unit;

[0104] S502: The address translation unit receives the first instruction transmitted by the instruction processor, converts the virtual address carried in the first instruction into a physical address, and obtains a second instruction;

[0105] S503: The target execution unit executes the second instruction to obtain an instruction execution result.

[0106] In an embodiment of the present disclosure, the instruction processing device obtains a first instruction and transmits the first instruction to the address translation unit; after receiving the first instruction, the address translation unit converts the virtual address in the first instruction into a physical address to obtain a second instruction; the target execution unit executes the second instruction to obtain an instruction execution result, so that before the instruction is distributed to a specific execution unit, the address translation unit is used to convert the virtual address carried in the instruction into a physical address; when the instruction is distributed to the execution unit, the virtual address carried in the instruction has been translated into a physical address by the address translation unit, so that the execution unit does not need to perform the address translation process, and thus there is no need to set up a TLB and a PTW in the execution unit, reducing the volume of the execution unit; at the same time, there will be no global memory bandwidth problem caused by multiple execution units performing concurrent page table queries, reducing the bandwidth pressure on the global memory.

[0107] In a possible implementation manner, it further includes:

[0108] The address translation unit transmits the second instruction to the instruction distribution unit;

[0109] After receiving the second instruction transmitted by the address translation unit, the instruction distribution unit determines the target execution unit among multiple execution units for the second instruction and transmits the second instruction to the target execution unit.

[0110] In a possible implementation manner, the target execution unit executes the second instruction, including:

[0111] The target execution unit parses the physical address carried in the second instruction and accesses the physical address.

[0112] In a possible implementation manner, the address translation unit receives the first instruction transmitted by the instruction processor, converts the virtual address carried in the first instruction into a physical address, and obtains a second instruction, including:

[0113] The address translation unit parses the virtual address from the first instruction, determines the physical address corresponding to the virtual address, and generates a second instruction based on the physical address and the first instruction.

[0114] In a possible implementation, it further includes: the mapping relationship between the virtual address and the physical address stored in the address translation query cache unit, and in the case where the mapping relationship corresponding to the virtual address does not exist in the cache unit, obtaining the physical address corresponding to the virtual address from the memory through page table query.

[0115] In a possible implementation, it further includes:

[0116] The address translation unit stores the mapping relationship between the physical address obtained from the memory and the virtual address into the cache unit.

[0117] In a possible implementation, the address translation unit converts the virtual address carried in the first instruction into a physical address to obtain a second instruction, including:

[0118] Changing the target flag bit in the first instruction from a first value to a second value, and replacing the virtual address with the physical address to obtain the second instruction;

[0119] Wherein, the first value indicates that the address carried in the instruction is a virtual address; the second value indicates that the address carried in the instruction is a physical address.

[0120] The description of the processing flow of each step in the method and the interaction flow between each step can refer to the relevant description in the foregoing instruction processing device embodiment, and will not be elaborated here.

[0121] The embodiments of the present disclosure further provide a computer program, which when executed by a processor implements any one of the methods in the foregoing embodiments. The computer program product can be specifically implemented in a manner of hardware, software, or a combination thereof. In an alternative embodiment, the computer program product is specifically embodied as a computer storage medium. In another alternative embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.

[0122] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein. In several embodiments provided in the present disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical functional division, and there can be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings, direct couplings, or communication connections shown or discussed with each other can be through some communication interfaces. The indirect couplings or communication connections between devices or units can be electrical, mechanical, or other forms.

[0123] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0124] In addition, in each embodiment of the present disclosure, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0125] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present disclosure. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0126] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present disclosure, used to illustrate the technical solutions of the present disclosure, rather than limiting them. The protection scope of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the art within the technical scope disclosed by the present disclosure can still modify the technical solutions described in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should all be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.

Claims

1. An instruction processing device, characterized in that, Comprising: An instruction processor, an address translation unit, and a target execution unit; The instruction processor is configured to obtain a first instruction and transmit the first instruction to the address translation unit; the first instruction includes an address and a flag bit corresponding to the address; The address translation unit is configured to receive the first instruction transmitted by the instruction processor, convert the virtual address carried in the first instruction into a physical address, and obtain a second instruction; The target execution unit is configured to execute the second instruction to obtain an instruction execution result; The address translation unit includes: an instruction parsing subunit, an address conversion subunit, and an instruction conversion subunit; Wherein, the instruction parsing subunit is connected to the instruction processor and is configured to parse the virtual address from the first instruction; The address conversion subunit is configured to determine the physical address corresponding to the virtual address parsed by the instruction parsing subunit; The instruction conversion subunit is configured to generate a second instruction based on the physical address determined by the address conversion subunit and the first instruction; The instruction parsing subunit includes: a first circuit; the first circuit is configured to perform a logical operation on the flag bit to generate a control signal; the control signal indicates that when the flag bit is a first value, the data of the address bit corresponding to the flag bit is transmitted to the address conversion subunit; the first value indicates that the address carried in the first instruction is a virtual address; The address conversion subunit includes: a third circuit; the third circuit is configured to receive the physical address transmitted by the address conversion subunit, and through logical operation, change the flag bit from the first value to a second value, and generate the second instruction based on the flag bit with the value changed and the physical address.

2. The instruction processing device according to claim 1, wherein Further comprising: An instruction distribution unit; The address translation unit is further configured to transmit the second instruction to the instruction distribution unit; The instruction distribution unit is configured to, after receiving the second instruction transmitted by the address translation unit, determine the target execution unit among multiple execution units for the second instruction, and transmit the second instruction to the target execution unit.

3. The instruction processing device according to claim 1 or 2, characterized in that, When executing the second instruction, the target execution unit is configured to: Parse the physical address carried in the second instruction and access the physical address.

4. The instruction processing device according to claim 1 or 2, characterized in that, The address translation unit includes: a cache unit; The address translation unit is configured to query the mapping relationship between the virtual address and the physical address stored in the cache unit, and in the case where there is no mapping relationship corresponding to the virtual address in the cache unit, obtain the physical address corresponding to the virtual address from the memory through page table query.

5. The instruction processing device according to claim 4, characterized in that, The address translation unit is further configured to store the mapping relationship between the physical address obtained from the memory and the virtual address into the cache unit.

6. The instruction processing device according to claim 1 or 2, characterized in that, When converting the virtual address carried in the first instruction into a physical address to obtain a second instruction, the address translation unit is configured to: Change the target flag bit in the first instruction from the first value to the second value, and replace the virtual address with the physical address to obtain the second instruction; Among them, the first value indicates that the address carried in the instruction is a virtual address; the second value indicates that the address carried in the instruction is a physical address.

7. An instruction processing method, characterized in that, Include; The instruction processor obtains a first instruction and transmits the first instruction to the address translation unit; the first instruction includes an address and a flag bit corresponding to the address; The address translation unit receives the first instruction transmitted by the instruction processor, converts the virtual address carried in the first instruction into a physical address, and obtains a second instruction; The target execution unit executes the second instruction to obtain an instruction execution result; The address translation unit includes: an instruction parsing subunit, an address conversion subunit, and an instruction conversion subunit; Among them, the instruction parsing subunit is connected to the instruction processor and is used to parse the virtual address from the first instruction; The address conversion subunit is used to determine the physical address corresponding to the virtual address parsed by the instruction parsing subunit; The instruction conversion subunit is used to generate a second instruction based on the physical address determined by the address conversion subunit and the first instruction; The instruction parsing subunit includes: a first circuit; the first circuit is used to perform a logical operation on the flag bit to generate a control signal; the control signal indicates that when the flag bit is the first value, the data of the address bit corresponding to the flag bit is transmitted to the address conversion subunit; the first value indicates that the address carried in the first instruction is a virtual address; The address conversion subunit includes: a third circuit; the third circuit is used to receive the physical address transmitted by the address conversion subunit, and through logical operation, change the flag bit from the first value to the second value, and generate the second instruction based on the flag bit with the value changed and the physical address.

8. The instruction processing method according to claim 7, wherein Also include: The address translation unit transmits the second instruction to the instruction distribution unit; After receiving the second instruction transmitted by the address translation unit, the instruction distribution unit determines the target execution unit among multiple execution units for the second instruction, and transmits the second instruction to the target execution unit.

9. The instruction processing method according to claim 7 or 8, characterized in that, The execution of the second instruction by the target execution unit includes: The target execution unit parses the physical address carried in the second instruction and accesses the physical address.

10. The instruction processing method according to claim 7 or 8, characterized in that, The address translation unit receives the first instruction transmitted by the instruction processor, converts the virtual address carried in the first instruction into a physical address, and obtains a second instruction, including: The address translation unit parses the virtual address from the first instruction, determines the physical address corresponding to the virtual address, and generates a second instruction based on the physical address and the first instruction.

11. The instruction processing method according to claim 7 or 8, wherein Also include: The mapping relationship between the virtual address and the physical address stored in the address translation query cache unit, and in the case where there is no mapping relationship corresponding to the virtual address in the cache unit, obtaining the physical address corresponding to the virtual address from the memory through page table query.

12. The instruction processing method according to claim 11, wherein Also include: The address translation unit stores the mapping relationship between the physical address obtained from the memory and the virtual address into the cache unit.

13. The instruction processing method according to claim 7 or 8, characterized in that The address translation unit converts the virtual address carried in the first instruction into a physical address to obtain a second instruction, including: changing the target flag bit in the first instruction from a first value to a second value, and replacing the virtual address with the physical address to obtain the second instruction; wherein, the first value indicates that the address carried in the instruction is a virtual address; the second value indicates that the address carried in the instruction is a physical address.

Citation Information

Patent Citations

  • Real-time hardware-assisted GPU tuning using machine learning

    US20190213775A1

  • Translation of virtual addresses in a computer graphics system

    US5313577A