Instruction operation method and device, medium and computer program product
By pre-allocating the target instruction association data of the processor in the electronic device multiple times to determine the optimal allocation result for cache space utilization, the problem of low cache utilization is solved and the efficiency of processor execution instructions is improved.
Patent Information
- Application Number
- CN202510281255.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-10
AI Technical Summary
The processor in electronic devices cannot effectively store more data required to execute instructions into the cache, resulting in low cache utilization and affecting the efficiency of the processor's execution of instructions.
By preallocating M times of N associated data of the target instruction to be run in off-chip memory and the storage space in the cache, a first preallocated result with the largest sum of the accumulated space occupancy of each associated data in the cache, and running the target instruction based on this result.
It effectively improves the utilization rate of cache, reduces the frequency of the processor reading data from off-chip memory, and thus improves the efficiency of the processor to execute target instructions.
Smart Images

Figure CN120123263A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technologies, and in particular, to an instruction execution method, device, medium, and computer program product. Background Art
[0002] During the use of an electronic device, the electronic device can read and / or write data from / to a main memory through a processor, and the processor performs arithmetic processing on the data, enabling the normal operation of the electronic device. However, since the arithmetic speed of the processor is much higher than the read / write speed of the main memory, in some cases, a cache may be configured in the electronic device. The electronic device can store some data in the cache through the processor. When the processor needs to execute an instruction based on this part of the data, the processor can quickly read the data from the cache, thereby improving the efficiency of the processor in executing instructions.
[0003] However, the storage space of the cache is limited. During the process of the electronic device storing data in the cache through the processor, if data is randomly stored in the cache, the utilization rate of the cache will be reduced. For example, when data for an instruction with a low execution frequency is stored in the cache, there may not be enough storage space in the cache to store data for an instruction with a high execution frequency. In this way, the processor needs to read the data required to execute the instruction from the main memory, affecting the execution efficiency of the instruction based on this part of the data. Summary of the Invention
[0004] To solve the problem that an electronic device cannot store more data required to execute instructions in the cache through a processor, this application provides an instruction execution method, device, medium, and computer program product.
[0005] In a first aspect, an embodiment of this application provides an instruction execution method, which is applied to an electronic device. The electronic device includes a processor and an off-chip memory outside the processor. The processor includes a cache. Moreover, the method includes: determining N associated data of a target instruction to be executed; performing M pre-allocations on the storage spaces of the N associated data in the off-chip memory and the cache to obtain M pre-allocation results, where each pre-allocation result includes the storage spaces corresponding to the N associated data, and for each pre-allocation result, the storage spaces of at least some of the N associated data are located in the cache, and M is a positive integer greater than 1; determining a first pre-allocation result among the M pre-allocation results, where the sum of the cumulative space occupancies of the respective associated data whose storage spaces are located in the cache is the largest, and the cumulative space occupancy indicates the product of the storage space occupied by the corresponding associated data and the number of times the corresponding associated data is read during the execution of the target instruction; and executing the target instruction based on the first pre-allocation result corresponding to the N associated data.
[0006] In some alternative embodiments, the target instruction may refer to an instruction for the electronic device to execute a game through the processor, an instruction for rendering or processing a game interface to be displayed; or in an autonomous driving scenario, the target instruction may refer to instructions such as pedestrian detection, obstacle detection, and emergency avoidance executed by the electronic device through the processor; or in a scenario of running a search engine, the target instruction may refer to a search instruction executed by the electronic device through the processor.
[0007] In some alternative embodiments, the associated data of the target instruction may include, but is not limited to: input data required by the processor during the execution of the target instruction by the electronic device through the processor, intermediate data generated during the execution of the target instruction, and output data corresponding to the execution result of the target instruction.
[0008] In some alternative embodiments, the cache may be a storage unit in the processor for providing storage space to the processor. The off-chip memory may be a memory disposed outside the processor, such as the memory or hard disk of the electronic device.
[0009] In some alternative embodiments, since the total data volume of the N associated data of the target instruction is constant, the larger the sum of the cumulative space occupancies of the respective associated data located in the cache, the less the amount of associated data read from the off-chip memory when the electronic device runs the target instruction through the processor, thereby improving the utilization rate of the cache.
[0010] Therefore, in the present application, by determining the first pre-allocation result with the largest sum of cumulative space occupancies among multiple pre-allocation results and running the target instruction based on the first pre-allocation result corresponding to the N associated data, the utilization rate of the cache can be effectively improved, the frequency of the processor reading data from the off-chip memory can be reduced, and it is beneficial to improve the efficiency of the processor executing the target instruction.
[0011] In a possible implementation of the first aspect above, performing M times of pre-allocation on the storage spaces of the N associated data in the off-chip memory and the cache includes: performing M times of pre-allocation on the storage spaces of the N associated data in the off-chip memory and the cache based on the cumulative space occupancies respectively corresponding to the N associated data.
[0012] In some alternative embodiments, the electronic device may set an allocation sequence of the N associated data, and when performing different pre-allocations, the order of the N associated data in the allocation sequence may be different. During the process of pre-allocating storage space for the N associated data each time, the electronic device may sequentially allocate the storage spaces in the off-chip memory and the cache to the N associated data in the order from the front to the back of the allocation sequence. For example, determining the allocation sequence of the N associated data based on the cumulative space occupancies respectively corresponding to the N associated data.
[0013] Therefore, corresponding to the allocation sequence of N associated data determined by the cumulative space occupation respectively corresponding to the N associated data, the electronic device, through the processor, pre-allocates the storage spaces of the N associated data in the off-chip memory and the cache M times in sequence according to the allocation sequence. The greater the sum of the cumulative space occupations of at least some of the associated data allocated to the cache among the N associated data, the greater the possibility. Thus, a better pre-allocation result can be obtained with a smaller number of pre-allocations, and the computational amount of the electronic device for pre-allocation through the processor can be reduced.
[0014] In a possible implementation of the first aspect above, pre-allocating the storage spaces of the N associated data in the off-chip memory and the cache M times includes: in the first pre-allocation among the M pre-allocations: initializing the allocation sequence based on the first order of the N associated data; sequentially allocating storage spaces to the N associated data in the allocation sequence, where:
[0015] corresponding to the i-th associated data in the allocation sequence satisfying the first condition, allocating the storage space in the cache to the i-th associated data, where i is a positive integer greater than or equal to 1 and less than or equal to N, or,
[0016] corresponding to the i-th associated data not satisfying the first condition, allocating the storage space in the off-chip memory to the i-th associated data and adjusting the i-th associated data to the tail of the allocation sequence.
[0017] In some optional embodiments, during the allocation process of one pre-allocation by the electronic device through the processor, the position of the associated data in the allocation sequence can be adjusted based on whether the associated data is allocated in the cache. Thus, during the allocation process of the next pre-allocation, the electronic device, through the processor, can allocate the storage spaces of the N associated data based on the order obtained by adjusting the allocation sequence last time.
[0018] In a possible implementation of the first aspect above, the first condition includes: the free space in the cache is greater than or equal to the storage space required by the associated data for which the storage space is to be allocated; or, the associated data stored in the first storage space in the cache has a storage space reuse relationship with the associated data for which the storage space is to be allocated;
[0019] wherein, the first storage space is greater than or equal to the storage space required by the associated data for which the storage space is to be allocated, or the sum of the first storage space and the adjacent free storage space after the first storage space is greater than or equal to the storage space required by the associated data for which the storage space is to be allocated, and there is no free space adjacent to the first storage space in front of the first storage space.
[0020] In some alternative embodiments, for example, when the free space in the cache is greater than the storage space required for the i-th associated data; or, when the first associated data stored in the first storage space in the cache has a storage space reuse relationship with the i-th associated data, the i-th associated data satisfies the first condition.
[0021] In a possible implementation of the first aspect above, that the associated data to be stored in the first storage space in the cache has a storage space reuse relationship with the associated data for which storage space is to be allocated includes: the starting address of the associated data for which storage space is to be allocated in the cache is the same as the starting address of the associated data stored in the first storage space in the cache.
[0022] In some alternative embodiments, having a storage space reuse relationship may mean that: the starting address of the associated data for which storage space is to be allocated in the cache is the same as the starting address of the associated data stored in the first storage space in the cache. For example, if data 2 and data 1 have a storage space reuse relationship, the starting address of data 2 in the cache can be configured to be the same as the starting address of data 1 in the cache.
[0023] In some alternative embodiments, when the electronic device stores each associated data in the cache through the processor, data A and data B have a storage space reuse relationship. When data A is stored in a certain storage space, after the electronic device generates or reads data B through the processor, data B can be stored in the storage space where data A is stored to overwrite the storage space occupied by data A; when data B is stored in a certain storage space, after the electronic device generates or reads data A through the processor, data A can be stored in the storage space where data B is stored to overwrite the storage space occupied by data B.
[0024] In a possible implementation of the first aspect above, in the j-th pre-allocation among M times of pre-allocation: storage spaces are sequentially allocated to N associated data in the allocation sequence after the (j - 1)-th pre-allocation, where j is a positive integer greater than 1 and less than or equal to M, and:
[0025] If the h-th associated data in the allocation sequence after the (j - 1)-th pre-allocation satisfies the first condition, the storage space in the cache is allocated to the h-th associated data, where h is a positive integer greater than or equal to 1 and less than or equal to N, or,
[0026] If the h-th associated data does not satisfy the first condition, the storage space in the off-chip memory is allocated to the h-th associated data, and the h-th associated data is adjusted to the end of the allocation sequence after the (j - 1)-th pre-allocation.
[0027] In a possible implementation of the above first aspect, the first order is the order from largest to smallest according to the cumulative space occupied by the N associated data.
[0028] In some alternative embodiments, the first order can also be determined by the electronic device through the processor according to the number of associated data that can have a storage space multiplexing relationship with each associated data among the N associated data in descending order; or it can be determined by the electronic device through the processor according to the storage space corresponding to the N associated data in descending order; or, it can be determined by the order randomly assigned by the electronic device through the processor.
[0029] In a possible implementation of the above first aspect, the processor includes at least one of the following: a central processing unit, a graphics processing unit, a neural network processing unit, a digital signal processing unit, and a field programmable gate array.
[0030] In a second aspect, an embodiment of the present application provides an electronic device, which includes: a memory for storing instructions executed by one or more processors of the electronic device; and a processor, which is one of the processors of the electronic device, for executing the instructions stored in the electronic device to implement the above first aspect and the instruction running method mentioned in the above first aspect.
[0031] In a third aspect, the present application provides a readable storage medium, on which instructions are stored, and when the instructions are executed on the electronic device, the electronic device is caused to execute the first aspect of the present application and the instruction running method mentioned in the above first aspect.
[0032] In a fourth aspect, an embodiment of the present application provides a computer program product, including: computer programs / instructions, and the computer programs / instructions are executed by the processor to implement the computer program code of the above first aspect and the instruction running method mentioned in the above first aspect.
[0033] For the beneficial effects of the above second aspect, third aspect, and fourth aspect, reference can be made to the above first aspect and the relevant descriptions in various possible implementations of the first aspect, and details are not elaborated here. Description of the Drawings
[0034] Figure 1 According to some embodiments of the present application, a schematic diagram of a data reading and storage scenario is shown;
[0035] Figure 2 According to some embodiments of the present application, a schematic diagram of writing data into a cache is shown;
[0036] Figure 3 According to some embodiments of the present application, a schematic diagram of an instruction running process is shown;
[0037] Figure 4 According to some embodiments of the present application, a schematic diagram of the first pre - allocation of N associated data is shown;
[0038] Figure 5A According to some embodiments of the present application, a schematic diagram of allocating storage space for N associated data in the first pre - allocation is shown;
[0039] Figure 5B According to some embodiments of the present application, another schematic diagram of allocating storage space for N associated data in the first pre - allocation is shown;
[0040] Figure 6 According to some embodiments of the present application, yet another schematic diagram of allocating storage space for N associated data in the first pre - allocation is shown;
[0041] Figure 7 According to some embodiments of the present application, a schematic diagram of the second pre - allocation of N associated data is shown;
[0042] Figure 8 According to some embodiments of the present application, a schematic diagram of allocating storage space for N associated data in the second pre - allocation is shown;
[0043] Figure 9 According to some embodiments of the present application, a schematic diagram of the j - th pre - allocation of N associated data is shown;
[0044] Figure 10 According to some embodiments of the present application, a block diagram of an electronic device is shown. Detailed implementation manners
[0045] The illustrative embodiments of the present application include but are not limited to an instruction execution method, a device, a medium, and a computer program product.
[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application.
[0047] The electronic devices mentioned in this application will be introduced below. It can be understood that the method provided by the embodiments of this application can be applied to any electronic device, including but not limited to mobile stations (MS), mobile terminals (MT), etc. For example, the electronic device can be a mobile phone, smart TV, wearable device, tablet computer (Pad), desktop computer, laptop computer, virtual reality (VR) device, augmented reality (AR) device, terminal in industrial control, terminal in self-driving, terminal in remote medical surgery, terminal in smart grid, terminal in transportation safety, terminal in smart city, terminal in smart home, and so on. The embodiments of this application do not limit the specific form of the electronic device.
[0048] To facilitate the understanding of the technical solution of this application, the related terms will be introduced below.
[0049] (1) Cache, a storage unit with a read and write speed higher than that of the main memory in an electronic device, such as static random-access memory (SRAM), etc. In some embodiments, since the operation speed of the processor in the electronic device is much higher than the read and write speed of the main memory, such as processors like central processing unit (CPU), neural processing unit (NPU), etc., during the process of the electronic device executing instructions through the processor, if the processor reads data from the main memory, the execution efficiency of the processor will be affected because of the slow speed of waiting for the data to be transmitted from the main memory to the processor. In response to this situation, the processor can store the frequently read and / or written data in the cache. When the processor needs this data, it can quickly read it from the cache, which can improve the execution efficiency of the processor. In the embodiments of this application, the cache and the buffer can also be referred to as global memory (GM).
[0050] (2) Bandwidth, which refers to the rate of data transmission between the main memory and the processor. The higher the bandwidth, the faster the processor can read and write data from the main memory. The more times the processor reads and writes data from the main memory, the greater the bandwidth pressure.
[0051] The following describes the situation of reading and writing data during the process of an electronic device executing instructions through a processor with reference to the accompanying drawings.
[0052] In some alternative embodiments, during the process of an electronic device executing instructions through a processor, the electronic device needs to read data from the main memory through the processor to execute one or more instructions, and store the execution result of the instructions into the main memory so that the electronic device can read the execution result through the processor during the process of executing other instructions. To ensure the execution efficiency of the processor, the processor can store the execution result into the cache so that the electronic device can quickly read data from the cache through the processor when executing instructions with the execution result as input data. However, the storage space of the cache is limited. If the electronic device arbitrarily stores data such as execution results into the cache through the processor, it will cause the storage space of the cache to be insufficient.
[0053] Exemplarily, as Figure 1 shown in the schematic diagram of the data reading and storage scenario. As Figure 1 shown, the electronic device may include a processor 101 and an off-chip memory 104. Among them, the processor 101 may include a computing unit 102 and a cache 103.
[0054] It should be noted that the processor 101 may be a graphics processing unit (GPU), a neural processing unit (NPU), a digital signal processor (DSP), a field-programmable gate array (FPGA), etc.
[0055] The cache 103 may be a storage unit in the processor 101, and is used to provide storage space for the processor 101.
[0056] The off-chip memory 104 may be a memory disposed outside the processor 101, such as the memory or hard disk of the electronic device.
[0057] When the electronic device executes an instruction through the computing unit 102 of the processor 101, the computing unit 102 may store the execution result of the instruction in the cache 103. Since the storage space of the cache is limited, during the process of the computing unit 102 storing the execution result in the cache 103, in the case where the execution result cannot be stored in the cache 103 due to insufficient storage space of the cache 103, etc., the computing unit 102 may store the execution result in the off-chip memory 104. Thus, when the computing unit 102 executes an instruction with the execution result as the input data, it needs to read data from the off-chip memory 104, which affects the execution efficiency of the instruction with the execution result as the input data.
[0058] In some cases, before the electronic device executes a segment of instructions through the processor, the electronic device can obtain the storage space size (hereinafter referred to as the data size) required for each associated data of the segment of instructions (such as the input data required for the processor to execute the instruction, the intermediate data generated during the execution of the instruction, and the output data corresponding to the execution result of the segment of instructions) through the processor. Then, the electronic device can allocate storage space for each associated data in a way of offset assignment through the processor. Specifically, the electronic device can configure the corresponding storage space for the associated data according to the order of the data sizes of each associated data from large to small through the processor. For example, the electronic device can first allocate the storage space in the cache 103 to the associated data with the largest data size among each associated data according to the data size from large to small through the processor 101; when the cache 103 cannot store the associated data (such as insufficient storage space of the cache 103), then allocate storage space for the associated data in the off-chip memory 104, such as allocating storage space for the associated data in the double data rate (DDR) memory.
[0059] In some alternative embodiments, as Figure 2 shown in the schematic diagram of allocating storage space, before the electronic device executes an instruction through the processor 101, the associated data obtained by the electronic device through the processor 101 during the execution of the instruction by the processor 101 includes: Data 1, Data 2, Data 3... Data 10. Among them, Data 2 and Data 1 have a storage space reuse relationship, Data 3 and Data 1 have a storage space reuse relationship, Data 4 and Data 2 have a storage space reuse relationship, Data 5 and Data 3 have a storage space reuse relationship, Data 5 and Data 2 do not have a storage space reuse relationship, Data 6 and Data 4 do not have a storage space reuse relationship, and Data 6 and Data 5 do not have a storage space reuse relationship. And the data sizes corresponding to Data 1, Data 2, Data 3... Data 10 decrease in sequence.
[0060] Among them, two data with a storage space reuse relationship may refer to: when storing one of the two data, the one data can be stored in the storage space for storing the other data of the two data. That is to say, the one data and the other data do not need to be stored in the memory of the electronic device at the same time. For example, data A and data B have a storage space reuse relationship. When a certain storage space in the cache stores data A, after the electronic device generates or reads data B through the processor, data B can be stored in the storage space storing data A to overwrite the storage space occupied by data A; when a certain storage space in the cache stores data B, after the electronic device generates or reads data A through the processor, data A can be stored in the storage space storing data B to overwrite the storage space occupied by data B.
[0061] For example, that data 2 and data 1 have a storage space reuse relationship may mean that the starting address configuration of data 2 in the cache 103 can be the same as the starting address of data 1 in the cache 103. So that when the electronic device stores data 2 in the cache 103 through the processor 101, data 2 overwrites the storage space occupied by data 1. And when the storage space occupied by data 2 is smaller than the storage space occupied by data 1, the unoverwritten storage space generated when data 2 overwrites data 1 can be the free space of the cache. This free space is overwritten by the associated data having a storage space reuse relationship with data 1, such as data 3.
[0062] Moreover, data 3 and data 1 have a storage space reuse relationship. Corresponding to the electronic device needing data 2 and data 3 to execute instructions simultaneously, when the electronic device stores data 3 in the cache 103 through the processor 101, data 3 is stored continuously with data 2, and the termination address determined by the starting address and data size of data 2 is used as the starting address of data 3.
[0063] Furthermore, the electronic device can, through the processor 101, configure corresponding storage spaces for each associated data according to the order of the data sizes of each associated data from largest to smallest. The order corresponding to data 1, data 2, data 3... data 10 is the order in which the electronic device configures the storage spaces of the data through the processor 101.
[0064] According to the order of the data sizes of each associated data from largest to smallest, the electronic device, through the processor 101, first allocates the storage space in the cache 103 to data 1. For example, according to the data size of data 1, the storage address of data 1 in the cache 103 is [0X0000, 0X00FF].
[0065] Then, the electronic device sequentially allocates storage space in the cache 103 for data 2 and data 3 through the processor 101. Since data 2 and data 1 have a storage space reuse relationship, data 3 and data 1 have a storage space reuse relationship, and data 3 and data 2 can be continuously stored, according to the data sizes of data 2 and data 3, the storage address of data 2 in the cache 103 can be [0X0000, 0X00F0], and the storage address of data 3 in the cache 103 can be [0X00F0, 0X00FE].
[0066] Secondly, based on the fact that data 4 and data 2 have a storage space reuse relationship, data 5 and data 3 have a storage space reuse relationship, and data 5 and data 2 do not have a storage space reuse relationship, therefore, the electronic device sequentially allocates storage space in the cache 103 for data 4 and data 5 through the processor 101. The starting address of data 4 is the same as the starting address of data 2, and the starting address of data 5 is the same as the starting address of data 3.
[0067] Moreover, since data 4 and data 5 are not continuously stored, there is storage space for unallocated data between data 4 and data 5. Since data 6 and data 4 do not have a storage space reuse relationship, and data 6 and data 5 do not have a storage space reuse relationship, when the electronic device allocates storage space in the cache 103 for data 6 through the processor 101, in the case where the data memory size of data 6 is greater than the size of the storage space for unallocated data between data 4 and data 5, it is impossible to allocate storage space in the cache 103 for data 6, so that the electronic device needs to allocate storage space in the off-chip memory 104 for data 6 through the processor 101.
[0068] Thus, during the process of the electronic device executing instructions through the processor 101, the electronic device reads associated data during the instruction execution process based on the above allocation results through the processor 101, reads associated data such as data 1 - data 5 from the cache 103, and reads associated data such as data 6 from the off-chip memory 104. However, data 5 that occupies a large storage space does not necessarily mean that it is frequently read and written by the processor 101. If data 6 that occupies a small storage space is frequently read and written by the processor 101 and this data 6 is stored in the off-chip memory 104, the processor 101 needs to frequently read data 6 from the off-chip memory 104, resulting in the processor 101 spending a long time to read data 6 and affecting the execution efficiency of the instructions based on data 6.
[0069] Thus, in the case of determining the storage sequence only according to the storage space occupied by the data, without considering the number of times of reading and writing data when the electronic device executes instructions through the processor, it is caused that the frequently read and written data is not stored in the cache, which not only reduces the efficiency of the processor executing instructions, but also reduces the utilization rate of the cache.
[0070] In some alternative embodiments, the reading speeds of the data with a large occupied storage space and the data with a small occupied storage space but a high access frequency will both affect the execution efficiency of the instructions executed by the electronic device that rely on this part of the data. Based on this, the cumulative space occupancy of the data (which can also be referred to as the memory cumulative footprint, memory cumulative occupancy) can be used to represent the response degree of the reading speed of the data to the execution efficiency of the relevant instructions. The larger the cumulative space occupancy of the data, the larger the storage space occupied by the data and / or the higher the access frequency of the data.
[0071] Among them, the cumulative space occupancy can be determined by the data size of the data (which can also be referred to as a tensor) and the number of times of reading and writing the data. When the data sizes of the data are the same, the more the number of times of reading and writing the data, the larger the cumulative occupied space. When the number of times of reading and writing the data is the same, the larger the data size of the data, the larger the cumulative occupied space. For example, the cumulative occupied space of the data can be obtained by multiplying the data size of the data by the number of times of reading and writing the data; or the cumulative occupied space of the data can be obtained by a preset multiple (such as 0.1, 0.2) of the product of the data size of the data and the number of times of reading and writing the data.
[0072] To solve the above problems, the present application provides an instruction running method. In this method, before executing the target instruction, N associated data in the target instruction to be run are determined (for example, the input data required for the processor to execute the target instruction, the intermediate data generated during the execution of the target instruction, the output data corresponding to the execution result of the target instruction), and the cumulative space occupancy corresponding to each associated data. Then, the electronic device pre-allocates the storage spaces of the N associated data in the off-chip memory (such as a double data rate memory) and the cache (such as a cache) M times (such as 100 times) through the processor, and obtains M pre-allocation results. Wherein, each pre-allocation result includes the storage spaces corresponding to the N associated data, and the storage spaces of at least some of the N associated data in each pre-allocation result are located in the cache. Secondly, among the M pre-allocation results, the first pre-allocation result with the largest sum of the cumulative space occupancies of the associated data whose storage spaces are located in the cache is determined as the allocation result of the storage spaces of the N associated data. Finally, based on the storage space allocation result corresponding to the N associated data, the target instruction is run.
[0073] It should be noted that the total data volume of the N associated data of the target instruction remains unchanged. Therefore, the larger the sum of the cumulative space occupations of the respective associated data stored in the cache, the less data volume of the associated data read from the off-chip memory by the electronic device when running the target instruction through the processor, thereby improving the utilization rate of the cache.
[0074] In this way, by performing multiple pre-allocations on the N associated data and determining the first pre-allocation result with the largest sum of cumulative space occupations among the multiple pre-allocation results, the utilization rate of the cache can be effectively improved, the frequency of the processor reading data from the off-chip memory can be reduced, which is beneficial to improving the efficiency of the processor executing the target instruction.
[0075] It should be noted that the target instruction can refer to an instruction for rendering or processing a game interface to be displayed that the electronic device executes through the processor in a game scenario. Executing this instruction faster can prevent the electronic device from freezing; or it can refer to an instruction for inferring business data during data processing. Executing this instruction faster can obtain results faster and avoid business waiting for data; or it can refer to instructions for vehicle acceleration, emergency avoidance, pedestrian detection, etc. in an autonomous driving scenario. Executing this instruction faster can reduce the probability of accidents; or it can refer to an instruction for searching the input information and outputting results in a running search engine scenario. Executing this instruction faster can reduce the time for waiting for the output result. The embodiments of the present application do not limit the target instruction.
[0076] The following introduces the instruction running method mentioned in the embodiments of the present application with reference to the accompanying drawings.
[0077] Figure 3 The flowchart of an instruction running method mentioned in the present application is shown. It can be understood that the execution subject of these process steps can be the electronic device mentioned above, which will not be elaborated here. The process can include:
[0078] S301: Determine N associated data of the target instruction to be run.
[0079] In some optional embodiments, for example, in a game scenario, the target instruction can refer to an instruction for the electronic device to run a game through the processor, an instruction for rendering or processing a game interface to be displayed, etc.; for example, in an autonomous driving scenario, the target instruction can refer to instructions for pedestrian detection, obstacle detection, and emergency avoidance that the electronic device executes through the processor; and for example, in a running search engine scenario, the target instruction can refer to a search instruction that the electronic device executes through the processor. The embodiments of the present application do not limit the target instruction.
[0080] The associated data of the target instruction may include, but is not limited to: during the process of the electronic device executing the target instruction through the processor, the input data required by the processor, the intermediate data generated during the execution of the target instruction, and the output data corresponding to the execution result of the target instruction.
[0081] In this regard, through the processor, the electronic device can determine N pieces of associated data of the target instruction to be run.
[0082] S302: Perform M times of pre-allocation on the storage spaces of the N pieces of associated data in the off-chip memory and the cache to obtain M pre-allocation results.
[0083] It can be understood that the cache may refer to the cache 103 mentioned above, which can be a storage unit in the processor and is used to provide storage space for the processor. The off-chip memory may refer to the off-chip memory 104 mentioned above, such as the memory, hard disk, etc. of the electronic device.
[0084] In some optional embodiments, during the process of the electronic device pre-allocating storage space for the N pieces of associated data through the processor each time, the storage spaces corresponding to the N pieces of associated data can be allocated in different orders.
[0085] For example, the electronic device can set an allocation sequence of the N pieces of associated data, and when performing different pre-allocations, the order of the N pieces of associated data in the allocation sequence can be different. During the process of pre-allocating storage space for the N pieces of associated data each time, the electronic device can sequentially allocate the storage spaces in the off-chip memory and the cache to the N pieces of associated data from the front to the back of the allocation sequence. When the i-th associated data to be allocated meets the first condition, allocate the storage space in the cache to this associated data; when the i-th associated data to be allocated does not meet the first condition, allocate the storage space in the off-chip memory to this associated data.
[0086] In some optional embodiments, the first condition includes: the free space in the cache is greater than or equal to the storage space required by the associated data for which the storage space is to be allocated; or, the associated data stored in the first storage space in the cache has a storage space reuse relationship with the associated data for which the storage space is to be allocated; wherein, the first storage space is greater than or equal to the storage space required by the associated data for which the storage space is to be allocated, or the sum of the first storage space and the adjacent free storage space after the first storage space is greater than or equal to the storage space required by the associated data for which the storage space is to be allocated, and there is no free space adjacent to the first storage space in front of the first storage space.
[0087] For example, when the free space in the cache is greater than or equal to the storage space required for the i-th associated data; or when there is a storage space reuse relationship between the first associated data stored in the first storage space in the cache and the i-th associated data, the i-th associated data satisfies the first condition.
[0088] Moreover, having a storage space reuse relationship may mean that the starting address of the associated data for which storage space is to be allocated in the cache is the same as the starting address of the associated data stored in the first storage space in the cache.
[0089] In some alternative embodiments, during the allocation process of a single pre-allocation by the processor of the electronic device, the position of the associated data in the allocation sequence may be adjusted based on whether the associated data is allocated in the cache. In this way, during the allocation process of the next pre-allocation, the processor of the electronic device may allocate storage space to the N associated data based on the order obtained by adjusting the allocation sequence last time.
[0090] In some embodiments, for M times of pre-allocation of the storage space of N associated data in the off-chip memory and the cache, the first pre-allocation and the j-th pre-allocation of the storage space of N associated data in the off-chip memory and the cache for M times of pre-allocation may be performed based on the following S3021 and S3022, so as to obtain M pre-allocation results.
[0091] S3021: Perform the first pre-allocation of the storage space of N associated data in the off-chip memory and the cache for M times of pre-allocation to obtain the first pre-allocation result.
[0092] In some alternative embodiments, during the process of the first pre-allocation by the processor of the electronic device of the storage space of N associated data in the off-chip memory and the cache, the processor of the electronic device needs to allocate storage space to the N associated data based on the allocation sequence corresponding to the first order.
[0093] Among them, the first order may be determined by the processor of the electronic device in the order from largest to smallest of the cumulative space occupation corresponding to the N associated data. In addition, in some alternative embodiments, the first order may also be determined by the processor of the electronic device in the order from most to least of the number of associated data that can have a storage space reuse relationship with each associated data among the N associated data; or it may be determined by the processor of the electronic device in the order from largest to smallest of the corresponding storage space of the N associated data; or alternatively, it may be determined by the random allocation order of the processor of the electronic device. The embodiments of the present application do not limit the first order.
[0094] The electronic device can initialize an allocation sequence based on the first order of N associated data through the processor, and the N associated data in the allocation sequence are sorted according to the first order. Then, the electronic device allocates storage space to the N associated data in the allocation sequence one by one based on the allocation sequence corresponding to the first order of the N associated data through the processor.
[0095] If the i-th associated data corresponding to the allocation sequence meets the first condition, the storage space in the cache is allocated to the i-th associated data, where i is a positive integer greater than or equal to 1 and less than or equal to N; or, if the i-th associated data does not meet the first condition, the storage space in the off-chip memory is allocated to the i-th associated data, and the i-th associated data is adjusted to the end of the allocation sequence, obtaining a first pre-allocation result of M pre-allocations of the storage space of the N associated data in the off-chip memory and the cache, and obtaining an allocation sequence of the N associated data after the first pre-allocation, which is called the second order.
[0096] Exemplarily, for the above Figure 2 shown that the electronic device determines the first order according to the order of the corresponding storage space of each associated data from large to small through the processor. The electronic device first allocates storage space to Data 1 through the processor. Since the free space in the cache is greater than the storage space required by Data 1 at this time, Data 1 meets the first condition, and the storage space in the cache is allocated to Data 1, and the storage space is allocated to Data 1 from the starting address 0X0000 of the cache.
[0097] Secondly, according to the allocation sequence corresponding to the first order, the electronic device allocates storage space to Data 2 through the processor. At this time, the free space in the cache is less than the storage space required by Data 2, but Data 2 has a storage space reuse relationship with Data 1 in the first storage space. Therefore, Data 2 meets the first condition. The storage space in the cache is allocated to Data 2, and the storage space is allocated to Data 2 from the starting address 0X0000 of the cache.
[0098] Then, according to the allocation sequence corresponding to the first order, the electronic device allocates storage space to Data 3 through the processor. At this time, the free space in the cache is greater than the storage space required by Data 3, and Data 3 has a storage space reuse relationship with Data 1 in the first storage space. Therefore, Data 3 meets the first condition. The storage space in the cache is allocated to Data 3, and Data 3 and Data 2 are continuously stored in the cache. The termination address determined by the starting address and the data size of Data 2 is used as the starting address of Data 3.
[0099] Furthermore, according to the allocation sequence corresponding to the first order, the electronic device sequentially allocates storage space for Data 4 through the processor. At this time, the free space in the cache is less than the storage space required by Data 4, but Data 4 has a storage space reuse relationship with Data 2 in the first storage space. Therefore, Data 4 meets the first condition. The storage space in the cache is allocated to Data 4, and the storage space for Data 4 is allocated starting from the starting address 0X0000 of the cache.
[0100] When the electronic device allocates storage space for Data 5 through the processor according to the allocation sequence, the free space in the cache is greater than the storage space required by Data 5 at this time, but Data 5 does not have a storage space reuse relationship with Data 2 in the first storage space, and Data 5 has a storage space reuse relationship with Data 2 in the first storage space. If the storage space in the cache is allocated to Data 5, then the storage space for Data 5 needs to be allocated starting from the starting address of Data 3, so that Data 5 cannot be continuously stored with Data 4, resulting in free space between Data 4 and Data 5. Therefore, Data 5 does not meet the first condition, and the storage space in the off-chip memory is allocated to Data 5. Then, the storage space for Data 6 - Data 10 is allocated. The pre-allocation results of Data 1 - Data 4 in the cache can be referred to Figure 2 in the allocation results of Data 1 - Data 4.
[0101] In addition, when the electronic device allocates storage space for Data 6 through the processor according to the allocation sequence, the free space in the cache is greater than the storage space required by Data 6 at this time, but Data 6 does not have a storage space reuse relationship with Data 2 in the first storage space, and Data 6 cannot be continuously stored with Data 2. Therefore, Data 6 does not meet the first condition. The storage space in the off-chip memory is allocated to Data 6, and then the electronic device allocates storage space for Data 7 - Data 10 through the processor.
[0102] Therefore, in the process of allocating storage space for each associated data based on the first order, the associated data such as Data 1, Data 2, Data 3, Data 4 corresponding to the allocation meet the first condition, and the storage space in the cache is allocated to the associated data such as Data 1, Data 2, Data 3, Data 4; the associated data such as Data 5, Data 6 corresponding to the allocation do not meet the first condition, and the storage space in the off-chip memory is allocated to the associated data such as Data 5, Data 6.
[0103] Since the total data volume of the N associated data of the target instruction remains unchanged, the larger the sum of the cumulative space occupancies of the respective associated data stored in the cache, the smaller the amount of associated data read from the off-chip memory when the electronic device runs the target instruction through the processor. Therefore, in the embodiments of the present application, the electronic device, through the processor, determines a first order according to the cumulative space occupancies corresponding to the N associated data from largest to smallest, and performs a first pre-allocation on the N associated data based on the allocation sequence of the first order. In this way, the greater the likelihood that the sum of the cumulative space occupancies of at least some of the N associated data allocated to the cache is maximized. Thus, a better pre-allocation result can be obtained with fewer pre-allocation times, and the computational amount of pre-allocation by the electronic device through the processor can be reduced. Therefore, based on the cumulative space occupancies respectively corresponding to the N associated data, the electronic device, through the processor, can perform M pre-allocations on the storage spaces of the N associated data in the off-chip memory and the cache.
[0104] Exemplarily, as Figure 4 shown in the schematic diagram of the first pre-allocation. The associated data of the target instruction may include data 311, data 312, data 313, data 314, data 315, data 316, data 317, data 318, data 319, and data 320. And, as Figure 4 shown, the allocation sequence corresponding to the first order determined from largest to smallest based on the cumulative space occupancies of the multiple associated data is successively: data 311, data 312, data 313, data 314, data 315, data 316, data 317, data 318, data 319, data 320.
[0105] Furthermore, the electronic device, through the processor, based on the allocation sequence corresponding to the first order of each associated data, successively allocates storage spaces for each associated data in the allocation sequence. The electronic device, through the processor, first allocates a storage space for data 311. Since the free space in the cache is greater than the storage space required by data 311 at this time, data 311 meets the first condition, and the storage space in the cache is allocated to data 311, and the storage space in the cache is allocated to data 311 starting from the starting address 0X0000 of the cache. And according to the data size of data 311, the storage address of data 311 in the cache can be [0X0000, 0X00F0].
[0106] Then, according to the allocation sequence, the electronic device, through the processor, allocates a storage space for data 312. Since, in the case of allocating the storage space in the cache to data 311, the free space in the cache is greater than the storage space required by data 312, data 312 meets the first condition, and the storage space in the cache is allocated to data 312, and data 312 can be stored continuously with data 311 (as Figure 4The first area of the example 401). Also, since there is a storage space reuse relationship between data 311 and data 312, the storage space can be allocated to data 312 starting from the starting address 0X0000 of the cache (such as Figure 4 the second area of the example 402), and according to the data size of data 312, the storage address of data 312 in the cache can be [0X0000, 0X00E0].
[0107] In this regard, when choosing to allocate the storage space to data 312 starting from the starting address 0X0000 of the cache, when the electronic device executes the target instruction through the processor, data 312 can overwrite the storage space occupied by data 311, and the storage space corresponding to the storage address [0X00E0, 0X00F0] that is not overwritten by data 312 is free space. This free space is overwritten by the associated data that has a storage space reuse relationship with data 311. In this way, the occupancy of data in the cache can be reduced, which is beneficial to improving the utilization rate of the cache.
[0108] Similarly, according to the allocation sequence, the electronic device sequentially allocates storage space to data 313 through the processor, such as Figure 5A as shown, when data 312 and data 311 are stored continuously, the free space in the cache is the storage space between the termination address of data 312 and the termination address of the cache. At this time, the free space in the cache is smaller than the data size of data 313. Since there is a storage space reuse relationship between data 311 stored in the first storage space in the cache and data 313, data 313 meets the first condition. Allocate the storage space in the cache to data 313, and allocate the storage space to data 313 starting from the starting address 0X0000 of the cache (such as Figure 5A the third area of the example 403). Among them, the first storage space at this time can refer to the storage space corresponding to data 311 and data 312.
[0109] In some alternative embodiments, such as Figure 5B as shown, when allocating the storage space to data 312 starting from the starting address 0X0000 of the cache, the free space in the cache is the storage space between the termination address 0X00E0 of data 312 and the termination address of the cache. At this time, the free space in the cache is larger than the data size of data 313. And there is a storage space reuse relationship between data 311 stored in the first storage space in the cache and data 313, so data 313 meets the first condition. Allocate the storage space in the cache to data 313, and allocate the storage space to data 313 starting from the termination address 0X00E0 of data 312 (such as Figure 5B the fourth area of the example 404). For ease of description, the following takes the pre-allocation result shown in Figure 5A as an example to sequentially allocate storage space to each associated data.
[0110] Similarly, according to the allocation sequence, the electronic device, through the processor, further allocates storage space for the data 314, as Figure 6 shown, the free space in the cache is the storage space between the termination address of the data 313 and the termination address of the cache. At this time, the free space in the cache is larger than the data size of the data 314. However, since the data 314 does not have a storage space reuse relationship with the data 311 stored in the first storage space in the cache, and the data 314 has a storage space reuse relationship with the data 312 stored in the first storage space in the cache. If the storage space in the cache is allocated to the data 314, the start address of the data 314 will be the same as that of the data 312, which will result in free space between the data 314 and the data 313 (such as Figure 6 the fifth region 405 in the example). Therefore, the data 314 does not meet the first condition, and the data 314 is adjusted to the end of the first-order allocation sequence. In this way, according to the allocation sequence, the electronic device, through the processor, further allocates storage space for the data 315 - data 320 in turn, and the first pre-allocation result of the first allocation of M times of pre-allocation of the storage space of each associated data in the off-chip memory and the cache can be obtained, and the allocation sequence of each associated data after the first pre-allocation is obtained, which is called the second order.
[0111] After the first pre-allocation, at this time, the sum of the cumulative occupied spaces in the cache is determined by the cumulative occupied spaces corresponding to the data 311, data 312, data 313, data 315, data 316, and data 318 stored in the cache. And the allocation sequence after the first pre-allocation, that is, the second order is: data 311, data 312, data 313, data 315, data 316, data 318, data 314, data 317, data 319, data 320.
[0112] In some alternative embodiments, corresponding to Figures 4 - 6 each of the associated data shown, during the second pre-allocation of the storage space of each associated data in the off-chip memory and the cache by the electronic device through the processor, the electronic device, through the processor, based on the allocation sequence corresponding to the second order of each associated data, allocates storage space for each associated data in the allocation sequence in turn.
[0113] It can be understood that according to the first pre-allocation result, the data 311, data 312, data 313, data 315, data 316, and data 318 all meet the first condition. Therefore, to avoid repetition, the pre-allocation results of allocating the storage space in the cache to the data 311, data 312, data 313, data 315, data 316, and data 318 can refer to the above Figures 4 - 6The corresponding description will not be elaborated here. Therefore, when the electronic device allocates the storage space in the cache based on the allocation sequence corresponding to the second order of each associated data, it allocates the storage space in the cache to Data 311, Data 312, Data 313, Data 315, Data 316, and Data 318, and then allocates the storage space to Data 314.
[0114] For example Figure 7 As shown in the schematic diagram of the second pre-allocation, since the free space in the cache is less than the storage space required by Data 314 at this time, and the Data 316 stored in the first storage space in the cache does not have a storage space reuse relationship with Data 314, Data 314 does not meet the first condition, and the storage space in the off-chip memory is allocated to Data 314.
[0115] Similarly, according to the allocation sequence corresponding to the second order, the electronic device sequentially allocates the storage space to Data 317 through the processor, as Figure 8 shown. At this time, the free space in the cache is less than the storage space required by Data 317, but the Data 316 stored in the first storage space in the cache has a storage space reuse relationship with Data 317. Therefore, Data 317 meets the first condition, and the storage space in the cache is allocated to Data 317.
[0116] Then, according to the allocation sequence corresponding to the second order, the electronic device sequentially allocates the storage space to Data 319 and Data 320 through the processor. Corresponding to Data 314, Data 319, and Data 320 not meeting the first condition, Data 314, Data 319, and Data 320 are sequentially adjusted to the end of the second order, and the storage space in the off-chip memory is allocated to Data 314, Data 319, and Data 320.
[0117] After the second pre-allocation, at this time, the sum of the cumulative occupied spaces in the cache is determined by the cumulative occupied spaces corresponding to Data 311, Data 312, Data 313, Data 315, Data 316, Data 318, and Data 317 stored in the cache. And the allocation sequence after the second pre-allocation, that is, the second order is: Data 311, Data 312, Data 313, Data 315, Data 316, Data 318, Data 317, Data 314, Data 319, Data 320.
[0118] So far, the electronic device has performed two pre-allocations on the storage spaces of N associated data in the off-chip memory and the cache through the processor, and obtained two pre-allocation results.
[0119] In some optional embodiments, corresponding to the sum of the cumulative space occupancies corresponding to the second pre-allocation result being greater than the sum of the cumulative space occupancies of the first pre-allocation result, the second pre-allocation result can be used as the first pre-allocation result.
[0120] S3022: The j-th pre-allocation of the storage spaces of N associated data in the off-chip memory and the cache is performed M times to obtain the j-th pre-allocation result, where j is a positive integer greater than 1 and less than or equal to M.
[0121] In some alternative embodiments, referring to the processes of the first and second pre-allocations of each associated data in S3021, during the process of the electronic device performing the j-th pre-allocation of the storage spaces of N associated data in the off-chip memory and the cache through the processor, the electronic device, through the processor, based on the allocation sequence corresponding to the j-th order of the N associated data, sequentially allocates storage spaces to the N associated data in the allocation sequence corresponding to the j-th order.
[0122] If the h-th associated data corresponding to the allocation sequence corresponding to the j-th order (i.e., the allocation sequence after the (j - 1)-th pre-allocation) satisfies the first condition, the storage space in the cache is allocated to the h-th associated data, where h is a positive integer greater than or equal to 1 and less than or equal to N; or, if the h-th associated data does not satisfy the first condition, the storage space in the off-chip memory is allocated to the h-th associated data, and the h-th associated data is adjusted to the tail of the allocation sequence corresponding to the j-th order, obtaining one pre-allocation result of the j-th allocation of the storage spaces of N associated data in the off-chip memory and the cache, and obtaining the allocation sequence of the N associated data after the j-th pre-allocation.
[0123] Exemplarily, continue to refer to Figures 4 - 8 each of the associated data shown. For example, the allocation sequence corresponding to the j-th order is data 311, data 312, data 313, data 315, data 316, data 318, data 317, data 314, data 319, data 320. It can be understood that data 311, data 312, data 313, data 315, data 316, data 318, data 317 satisfy the first condition. To avoid repetition, the process of the electronic device performing pre-allocation on data 311, data 312, data 313, data 315, data 316, data 318, and data 317 through the processor is not elaborated here. And the pre-allocation results of data 311, data 312, data 313, data 315, data 316, data 318, data 317 in the cache can be referred to Figure 8 .
[0124] Furthermore, as Figure 9For the j-th pre-allocation shown, when the electronic device allocates storage space for data 314 through the processor according to the allocation sequence corresponding to the j-th order, since data 314 has a storage space multiplexing relationship with data 318 in the first storage space and the free space in the cache is greater than the storage space required by data 314, data 314 meets the first condition, and the storage space in the cache is allocated to data 314.
[0125] Secondly, when the electronic device allocates storage space for data 319 through the processor according to the allocation sequence corresponding to the j-th order, at this time the free space in the cache is greater than the storage space required by data 319, but since data 319 has no storage space multiplexing relationship with either data 317 or data 318 in the first storage space, data 319 does not meet the first condition, and the storage space in the off-chip memory is allocated to data 319.
[0126] Then, when the electronic device allocates storage space for data 320 through the processor according to the allocation sequence corresponding to the j-th order, at this time the free space in the cache is greater than the storage space required by data 320, but since data 320 has no storage space multiplexing relationship with either data 317 or data 318 in the first storage space, data 320 does not meet the first condition, and the storage space in the off-chip memory is allocated to data 320. In this way, storage space is allocated to each associated data with different allocation sequences, and when different pre-allocations are performed, the order of each associated data in the allocation sequence is different, obtaining multiple pre-allocation results.
[0127] Furthermore, by performing M pre-allocations on the storage space of N associated data in the off-chip memory and the cache, the electronic device can obtain M pre-allocation results through the processor. Among them, each pre-allocation result includes the storage space corresponding to N associated data, and the storage space of at least some of the N associated data in each pre-allocation result is located in the cache.
[0128] S303: Determine the first pre-allocation result with the largest sum of the cumulative space occupancies of each associated data whose storage space is located in the cache among the M pre-allocation results.
[0129] After obtaining the M pre-allocation results, the sum of the cumulative space occupancies of each associated data whose storage space is located in the cache (hereinafter referred to as the target value) can be determined for each pre-allocation result first. Then, the pre-allocation result with the largest target value can be determined as the first pre-allocation result. Thus, based on the first pre-allocation result corresponding to the N associated data, the target instruction is run.
[0130] Thus, by performing multiple pre-allocations on N associated data and determining the first pre-allocation result with the largest cumulative space occupancy sum of each associated data whose storage space is located in the cache among the multiple pre-allocation results, the utilization rate of the cache can be effectively improved, the frequency of the processor reading data from the off-chip memory can be reduced, which is beneficial to improving the efficiency of the processor executing the target instruction and alleviating the bandwidth pressure.
[0131] It can be understood that in the embodiments of the present application, the above instruction execution method can be executed by the electronic device 001. Figure 10 It is a block diagram of the electronic device 001 provided in the embodiments of the present application. In some embodiments, the electronic device 001 may include one or more processors 801, system control logic 805 connected to at least one of the processors 801, an off-chip memory 804 connected to the system control logic 805, and a network interface 807 connected to the system control logic 805.
[0132] In some embodiments, the processor 801 may include one or more single-core or multi-core processors. In some embodiments, the processor 801 may include any combination of a general-purpose processor and a dedicated processor (e.g., a graphics processor, an application processor, a baseband processor, etc.). In the embodiments where the electronic device 001 employs an evolved node B (eNB) or a radio access network (RAN) controller, the processor 801 may be configured to execute various corresponding embodiments. Among them, the processor 801 may correspond to the processor 101 mentioned above.
[0133] In some alternative embodiments, the processor 801 includes a cache 802, where the cache 802 may be a storage unit in the processor 801 for providing storage space for the processor 801. The cache 802 may include a level-1 cache (L1 Cache), a level-2 cache (L2 Cache), and a level-3 cache (L3 Cache). The processor 801 may store frequently read and / or written data in the cache 802. When the processor 801 needs these data, it can quickly read them from the cache 802, which can improve the execution efficiency of the processor 801.
[0134] In an embodiment of the present application, when the electronic device 001 executes the instruction 803 through the processor 801, the processor 801 may store at least some of the associated data that meets the first condition among the various associated data of the instruction 803 into the cache 802, so that the processor 801 can quickly read the associated data required to execute the instruction 803 from the cache 802. And store the associated data that does not meet the first condition among the various associated data of the instruction 803 into the off-chip memory 804. In this way, the utilization rate of the cache can be effectively improved, the frequency of the processor 801 reading data from the off-chip memory 804 can be reduced, and it is beneficial to improve the efficiency of the processor 801 executing the instruction 803.
[0135] Exemplarily, in an embodiment of the present application, taking the various associated data shown above Figure 4 as an example, during the first pre-allocation of each associated data by the electronic device 001 through the processor 801, the electronic device 001 allocates storage space for the data 311 through the processor 801. Since the free space in the cache 802 is greater than the storage space required by the data 311 at this time, the data 311 meets the first condition, and the storage space in the cache 802 is allocated to the data 311, and the storage space is allocated to the data 311 from the starting address 0X0000 of the cache 802. And according to the data size of the data 311, the storage address of the data 311 in the cache 802 can be [0X0000, 0X00F0].
[0136] Then, according to the allocation sequence, the electronic device 001 allocates storage space for the data 312 through the processor 801. Since the free space in the cache 802 is greater than the storage space required by the data 312 when the storage space in the cache 802 is allocated to the data 311, the data 312 meets the first condition, and the storage space in the cache 802 is allocated to the data 312. The data 312 can be stored continuously with the data 311 (such as the first region 401 in the above Figure 4 example). Also, since the data 311 and the data 312 have a storage space reuse relationship, the storage space can be allocated to the data 312 from the starting address 0X0000 of the cache 802 (such as the second region 402 in the above Figure 4 example), and according to the data size of the data 312, the storage address of the data 312 in the cache 802 can be [0X0000, 0X00E0].
[0137] In this regard, when the storage space for the data 312 is allocated starting from the starting address 0X0000 of the cache 802, when the electronic device 001 executes the instruction 803 through the processor 801, the data 312 can overwrite the storage space occupied by the data 311, and the storage space corresponding to the storage addresses [0X00E0, 0X00F0] that are not overwritten by the data 312 is free space. In this way, the occupancy of data in the cache 802 can be reduced, which is beneficial to improving the utilization rate of the cache 802.
[0138] Among them, the associated data of the instruction 803 can include, but is not limited to: during the process of the electronic device 001 executing the instruction 803 through the processor 801, the input data required by the processor 801, the intermediate data generated during the execution of the instruction 803, and the output data corresponding to the execution result of the instruction 803.
[0139] In some embodiments, the system control logic 805 may include any suitable interface controller to provide any suitable interface to at least one of the processors 801 and / or any suitable device or component communicating with the system control logic 805.
[0140] In some embodiments, the system control logic 805 may include one or more memory controllers to provide an interface to the off-chip memory 804. The off-chip memory 804 can be used to load and store data and / or instructions.
[0141] In some embodiments, the off-chip memory 804 of the electronic device 001 may include any suitable memory, such as a suitable dynamic random access memory (DRAM), hard disk, memory, etc.
[0142] The off-chip memory 804 includes: non-volatile memory and volatile memory. Among them, the non-volatile memory includes electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic random access memory (MROM), ferroelectric random access memory (FROM), and phase change memory (PCM). The volatile memory includes: static random access memory (SRAM), dynamic random access memory (DRAM), and pseudo-static random access memory (PSRAM).
[0143] The off-chip memory 804 may include: a temporary copy of the instruction 803. The instruction 803 may include: an instruction that causes the electronic device 001 to implement the instruction running method mentioned in the embodiments of the present application when executed by at least one of the processors 801, corresponding to the target instruction mentioned above. In some embodiments, the instruction 803, hardware, firmware, and / or its software components may alternatively be placed in the system control logic 805, the network interface 807, and / or the processor 801.
[0144] Exemplarily, in the embodiments of the present application, the above-mentioned Figure 9 each of the associated data shown is taken as an example. During the process of the electronic device 001 pre-distributing the j-th time to each of the associated data through the processor 801, according to the allocation sequence corresponding to the j-th order, when the electronic device 001 allocates storage space for the data 319 through the processor 801, at this time, the free space in the cache 802 is greater than the storage space required by the data 319. However, since the data 319 has no storage space reuse relationship with the data 317 and the data 318 in the first storage space, the data 319 does not meet the first condition, and the storage space in the off-chip memory 804 is allocated to the data 319.
[0145] Then, according to the allocation sequence corresponding to the j-th order, when the electronic device 001 allocates storage space for the data 320 through the processor 801, at this time, the free space in the cache 802 is greater than the storage space required by the data 320. However, since the data 320 has no storage space reuse relationship with the data 317 and the data 318 in the first storage space, the data 320 does not meet the first condition, and the storage space in the off-chip memory 804 is allocated to the data 320. Thus, by storing the associated data that does not meet the first condition in the off-chip memory 804, the probability that the associated data frequently read and written by the processor 801 is stored in the off-chip memory 804 can be reduced, thereby improving the efficiency of the processor 801 in executing the instruction 803.
[0146] The network interface 807 may include a transceiver for providing a radio interface for the electronic device 001, and further communicating with any other suitable device (such as a front-end module, an antenna, etc.) through one or more networks. In some embodiments, the network interface 807 may be integrated with other components of the electronic device 001. For example, the network interface 807 may be integrated with at least one of the processor 801, the off-chip memory 804, and a firmware device having instructions (not shown in the figure). When at least one of the processors 801 executes the instructions, the electronic device 001 implements the instruction running method mentioned in the embodiments of the present application.
[0147] The network interface 807 may further include any suitable hardware and / or firmware to provide a multiple-input multiple-output radio interface. For example, the network interface 807 may be a network adapter, a wireless network adapter, a telephone modem, and / or a wireless modem.
[0148] In some embodiments, at least one of the processors 801 may be packaged with the logic of one or more controllers for the system control logic 805 to form a system-in-package (SiP). In one embodiment, at least one of the processors 801 may be integrated with the logic of one or more controllers for the system control logic 805 on the same die to form a system-on-chip (SoC).
[0149] The electronic device 001 may further include: an input / output (I / O) device 806.
[0150] The instruction execution method provided by the embodiments of the present application (for example Figure 3 the instruction execution method mentioned in S301 - S303) may be executed by the processor 801 in the above-mentioned electronic device 001.
[0151] The embodiments of the present application provide an electronic device, including: a memory for storing instructions executed by one or more processors of the electronic device, and a processor, which is one of the one or more processors of the electronic device, for executing the instruction execution method provided by the embodiments of the present application.
[0152] The embodiments of the present application provide a computer-readable medium, on which instructions are stored, and when the instructions are executed on the electronic device 001, the electronic device 001 is caused to execute the instruction execution method provided by the embodiments of the present application.
[0153] The embodiments of the present application provide a computer program product, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the instruction execution method provided by the embodiments of the present application is implemented.
[0154] It should be noted that in the examples and the description of the present application, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one" does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0155] Although the present application has been illustrated and described by referring to certain embodiments of the present application, those of ordinary skill in the art should understand that various changes may be made in form and detail without departing from the scope of the present application.
Claims
1. A method for executing an instruction, characterized in that: Applied to an electronic device, the electronic device comprises a processor and an off-chip memory outside the processor, the processor comprises a cache; and the method comprises: Determine N associated data of the target instruction to be executed; Pre-allocating the storage space of the N associated data in the off-chip memory and the cache M times to obtain M pre-allocation results, wherein each of the pre-allocation results includes the storage space corresponding to the N associated data, and the storage space of at least part of the associated data among the N associated data in each of the pre-allocation results is located in the cache, and M is a positive integer greater than 1; Determine a first pre-allocation result, among the M pre-allocation results, in which the sum of the accumulated space occupied by each associated data in the cache is the largest, wherein the accumulated space occupancy indicates the product of the storage space occupied by the associated data corresponding to the indication and the number of times the associated data is read during the execution of the target instruction; The target instruction is executed based on the first pre-allocation results corresponding to the N associated data.
2. The method according to claim 1, characterized in that Pre-allocating storage space of the N associated data in the off-chip memory and the cache M times includes: Based on the accumulated space occupancy respectively corresponding to the N associated data, storage space of the N associated data in the off-chip memory and the cache is pre-allocated M times.
3. The method according to claim 1 or 2, characterized in that: The pre-allocating the storage space of the N associated data in the off-chip memory and the cache M times includes: In the first pre-allocation of the M pre-allocations: Initialize the allocation sequence based on the first order of the N associated data; Allocate storage space for the N associated data in the allocation sequence in sequence, wherein: The i-th associated data corresponding to the allocation sequence meets the first condition, and the storage space in the cache is allocated to the i-th associated data, where i is a positive integer greater than or equal to 1 and less than or equal to N, or, Corresponding to the i-th associated data not satisfying the first condition, a storage space in the off-chip memory is allocated to the i-th associated data, and the i-th associated data is adjusted to the end of the allocation sequence.
4. The method according to claim 3, characterized in that The first condition includes: The free space in the cache is greater than or equal to the storage space required by the associated data to which the storage space is to be allocated; Alternatively, the associated data stored in the first storage space in the cache has a storage space reuse relationship with the associated data of the storage space to be allocated; Among them, the first storage space is greater than or equal to the storage space required for the associated data of the storage space to be allocated, or the sum of the first storage space and the free storage space adjacent to the first storage space is greater than or equal to the storage space required for the associated data of the storage space to be allocated, and there is no free space adjacent to the first storage space in front of the first storage space.
5. The method according to claim 4, characterized in that The associated data stored in the first storage space in the cache and the associated data of the storage space to be allocated have a storage space reuse relationship, including: The starting address of the associated data to which the storage space is to be allocated in the cache is the same as the starting address of the associated data stored in the first storage space in the cache.
6. The method according to claim 4, characterized in that In the j-th pre-allocation of the M pre-allocations: Allocate storage space for the N associated data in the allocation sequence after the j-1th pre-allocation in sequence, where j is a positive integer greater than 1 and less than or equal to M, wherein: The h-th associated data in the allocation sequence after the j-1-th pre-allocation satisfies the first condition, and the storage space in the cache is allocated to the h-th associated data, where h is a positive integer greater than or equal to 1 and less than or equal to N, or, Corresponding to the hth associated data not satisfying the first condition, the storage space in the off-chip memory is allocated to the hth associated data, and the hth associated data is adjusted to the end of the allocation sequence after the j-1th pre-allocation.
7. The method according to claim 3, characterized in that The first order is an order from large to small according to the cumulative space occupation corresponding to the N associated data.
8. The method according to claim 1, characterized in that: The processor includes at least one of the following: a central processing unit, a graphics processor, a neural network processor, a digital signal processor, and a field programmable gate array.
9. An electronic device, characterized in that: It comprises: a memory for storing instructions executed by one or more processors of the electronic device, and the processor, which is one of the one or more processors of the electronic device, for executing the instruction execution method according to any one of claims 1 to 8.
10. A computer-readable medium, characterized in that The computer-readable medium stores instructions, which, when executed on a computer, enable the computer to execute the instruction execution method according to any one of claims 1 to 8.
11. A computer program product, characterized in that The invention comprises a computer program / instruction, which, when executed by a processor, implements the instruction execution method according to any one of claims 1 to 8.