Address allocation method and device for shared memory variable, electronic equipment and chip
By classifying and grouping shared memory variables in heterogeneous calculations, and generating interference graphs for address allocation, the problem of redundancy in address allocation of shared memory variables in the prior art is solved, and more efficient memory space utilization is achieved.
Patent Information
- Application Number
- CN202510027060.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-07
AI Technical Summary
In heterogeneous computing, the prior art has redundancy in the allocation of shared memory variable addresses, resulting in an increase in memory space consumption.
By classifying shared memory variables into local and global variables, and generating interference graphs based on the kernel function group corresponding to the global shared memory variables, grouping and address allocation, the address allocation of global shared memory variables is optimized.
It effectively reduces the redundancy of address allocation, optimizes the utilization of memory space, and reduces the consumption of memory space.
Smart Images

Figure CN119938329A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the fields of computer technology and intelligent computing, and in particular to a method, device, electronic device and chip for allocating addresses of shared memory variables. Background Art
[0002] In the field of heterogeneous computing, applications can be executed in parallel to improve computing performance. The pipeline control unit in a Single Instruction, Multiple Threads (SIMT) processor creates thread groups and schedules them for execution, during which all threads in the group execute the same instruction at the same time. In a specific processor, each group has 32 threads, corresponding to the 32 execution pipelines or channels in the SIMT processor. Typically, these thread groups can be organized into a large execution module called a thread block. Threads within a thread block generally share a block of computing and storage resources on the hardware, one of which is shared memory. All threads in a thread block can access the same shared memory.
[0003] Shared memory is a memory interface released by the hardware for user programming. Users can declare variables stored in shared memory in the application. The compiler and other software stacks will then allocate addresses to the variables and generate memory access instructions to read and write the variables. Shared memory is a powerful tool for writing well-optimized parallel programming code, which can significantly improve program execution performance. Summary of the invention
[0004] The present disclosure aims to solve one of the technical problems in the related art at least to some extent.
[0005] The first embodiment of the present disclosure provides a method for allocating addresses of shared memory variables, including:
[0006] Determine the local shared memory variable, the global shared memory variable and the kernel function group corresponding to each global shared memory variable according to the calling information of each shared memory variable;
[0007] Based on the kernel function group corresponding to each global shared memory variable, an interference graph corresponding to the global shared memory variable is generated, wherein the global shared memory variable is a node in the interference graph, and the connecting edge in the interference graph represents that the kernel function groups corresponding to two connected global shared memory variables respectively include at least one identical kernel function;
[0008] Based on the interference graph, the global shared memory variables are grouped, and the starting positions and alignment information corresponding to the global shared memory variables included in each group are determined;
[0009] According to the total size of all global shared memory variables and the size of the local shared memory variable, the termination address of the local shared memory variable is determined.
[0010] The second aspect of the present disclosure provides an address allocation device for a shared memory variable, including:
[0011] A first determination module is used to determine the local shared memory variable, the global shared memory variable and the kernel function group corresponding to each global shared memory variable according to the call information of each shared memory variable;
[0012] A generating module, configured to generate an interference graph corresponding to the global shared memory variable based on the kernel function group corresponding to each global shared memory variable, wherein the global shared memory variable is a node in the interference graph, and the connecting edge in the interference graph indicates that the kernel function groups corresponding to two connected global shared memory variables respectively include at least one identical kernel function;
[0013] A second determination module is used to group the global shared memory variables based on the interference graph, and determine the starting position and alignment information corresponding to the global shared memory variables included in each group;
[0014] The third determining module is used to determine the end address of the local shared memory variable according to the total size of all global shared memory variables and the size of the local shared memory variable.
[0015] The third aspect embodiment of the present disclosure proposes an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the address allocation method for shared memory variables proposed in the first aspect embodiment of the present disclosure is implemented.
[0016] The fourth aspect embodiment of the present disclosure proposes a chip, which includes a processing circuit and an interface circuit; wherein the interface circuit is used to obtain instructions and send the instructions to the processing circuit, and the processing circuit is used to execute the instructions to implement the address allocation method of shared memory variables proposed in the first aspect embodiment of the present disclosure.
[0017] The fifth aspect embodiment of the present disclosure proposes a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the address allocation method of the shared memory variable proposed in the first aspect embodiment of the present disclosure is implemented.
[0018] The method, device, electronic device and storage medium for allocating addresses of shared memory variables provided by the present disclosure have the following features:
[0019] Beneficial effects:
[0020] In the disclosed embodiment, all shared memory variables are first classified into local shared memory variables and global shared memory variables, and then the global shared memory variables are grouped according to the kernel function group corresponding to the global shared memory variables, and the starting position and alignment information of each group in the memory space are determined to complete the address allocation of the global shared memory variables. Afterwards, according to the total size of the global shared memory variables after group alignment and the size of the local shared memory variables, the end address of the local shared memory variables is determined to complete the address allocation of the local shared memory variables. Thereby, the address allocation of the global shared memory variables is optimized, the redundancy of the address allocation is effectively reduced, and the memory space consumption is reduced.
[0021] Additional aspects and advantages of the present disclosure will be given in part in the following description and in part will be obvious from the following description or learned through practice of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The above and / or additional aspects and advantages of the present disclosure will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0023] Figure 1 A schematic diagram of a flow chart of a method for allocating addresses of shared memory variables provided by an embodiment of the present disclosure;
[0024] Figure 2 A schematic diagram of an interference diagram provided in an embodiment of the present disclosure;
[0025] Figure 3 is a schematic diagram of an address allocation result provided by an embodiment of the present disclosure;
[0026] Figure 4a It is a schematic diagram of a kernel function calling a device-side function and a device-side function referencing a shared memory variable, as proposed in an embodiment of the present disclosure;
[0027] Figure 4b It is a schematic diagram of grouping global shared memory variables proposed in an embodiment of the present disclosure;
[0028] Figure 4c is a schematic diagram of allocation results corresponding to the address allocation method proposed in the present disclosure;
[0029] Figure 4d It is a schematic diagram of allocation results corresponding to the existing address allocation method;
[0030] Figure 5 is a flow chart of a method for allocating addresses of shared memory variables provided by another embodiment of the present disclosure;
[0031] Figure 6 A schematic diagram of a shared memory variable classification process provided by an embodiment of the present disclosure;
[0032] Figure 7 A schematic flow chart of a method for allocating addresses of shared memory variables provided by another embodiment of the present disclosure;
[0033] Figure 8 A schematic diagram of the structure of an address allocation device for shared memory variables provided by an embodiment of the present disclosure;
[0034] Fig. 9 A block diagram of an exemplary electronic device suitable for implementing the embodiments of the present disclosure is shown;
[0035] Fig.10 It is a schematic diagram of the structure of a chip proposed in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0036] Embodiments of the present disclosure are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present disclosure, and should not be construed as limiting the present disclosure.
[0037] The following describes the method, device, electronic device and storage medium for allocating addresses of shared memory variables according to embodiments of the present disclosure with reference to the accompanying drawings.
[0038] In the related art, the shared memory compiler allocates the shared memory variable address in the kernel function by collecting the shared memory variables called by multiple kernel functions, sorting all the shared memory variables based on the optimization algorithm, obtaining the address allocation result that is arranged as closely as possible, and performing the same address allocation in each kernel function. Although this method solves the problem of address consistency of the same shared memory variable in different kernel functions, there may be a situation where some kernel functions do not use some shared memory variables. At this time, allocating an address separately for all global shared memory variables will lead to redundancy in address allocation and increase the consumption of shared memory space.
[0039] Therefore, in the present disclosure, shared memory variables can be classified and grouped according to the calling conditions of shared memory variables in different kernel functions, and then addresses can be allocated, so that some shared memory variables can reuse the same address, optimizing the space allocation of shared memory variables and reducing space consumption.
[0040] Figure 1 A flowchart of a method for allocating addresses of shared memory variables provided in an embodiment of the present disclosure.
[0041] It should be noted that the address allocation method for shared memory variables of the embodiment of the present disclosure can be applied to the address allocation device for shared memory variables. In some possible embodiments, the device can be configured in an electronic device or chip so that the electronic device or chip can perform the function of allocating addresses to shared memory variables.
[0042] like Figure 1 As shown, the address allocation method of the shared memory variable may include the following steps:
[0043] Step 101: Determine, according to the call information of each shared memory variable, a local shared memory variable, a global shared memory variable, and a kernel function group corresponding to each global shared memory variable.
[0044] It is understandable that the entry function of a heterogeneous device program is usually composed of one or more called device-side functions (function, referred to as func). The entry function can be called a kernel function, which is responsible for executing parallel computing tasks on the device side. Multiple parallel threads in a device thread block can access shared memory through a unified kernel function. The kernel function, as the entry function of all programs, enjoys the hardware resource configuration of the thread block. Therefore, in a parallel programming kernel function, by declaring a shared memory variable, it can be indicated that when the shared memory variable is used in the program, it needs to be read or written from the shared memory.
[0045] It should be noted that the call of shared memory variables can be implemented in two ways. One is to directly declare the called shared memory variables in the kernel function, and the shared memory variables can be directly called by the kernel function. The other is to declare the called shared memory variables in the ordinary device-side function, and then the kernel function calls the shared memory variables by calling the ordinary device-side function. Therefore, the call information of each shared memory variable may include at least one of the following: directly calling one or more kernel functions of the shared memory variable, calling the device-side function of the shared memory variable, and calling one or more kernel functions of the device-side function. In the embodiment of the present disclosure, the compiler can be used to collect all the shared memory variables in the heterogeneous device-side program, analyze the calling relationship between these shared memory variables and the functions they are in and the kernel function, so as to determine the call information of each shared memory variable.
[0046] In the implementation of the present disclosure, all shared memory variables can be classified according to the calling information of each shared memory variable. Specifically, if a shared memory variable declaration is called by multiple kernel functions at the same time, then the shared memory variable can be classified as a global shared memory variable, that is, the global shared memory variable can be used to indicate a variable called by multiple kernel functions. Conversely, if a shared memory variable declaration is only called by a single kernel function, then the shared memory variable can be classified as a local shared memory variable, that is, the local shared memory variable can be used to indicate a variable called by a single kernel function. In this way, all shared memory variables can be classified, local shared memory variables and global shared memory variables can be determined, and all kernel functions called by each global shared memory variable constitute a kernel function group corresponding to the global shared memory variable.
[0047] For example, all shared memory variables in the heterogeneous device program include shared memory variables A, shared memory variables B, and shared memory variables C. Shared memory variables A and C are called by kernel functions kernel0 and kernel1 at the same time, and shared memory variable B is only called by kernel1. Then shared memory variables A and C are global shared memory variables, and shared memory variable B is a local shared memory variable.
[0048] Step 102: Generate an interference graph corresponding to each global shared memory variable based on the kernel function group corresponding to each global shared memory variable.
[0049] In the disclosed embodiment, the interference graph corresponding to the global shared memory variables is used to characterize whether the global shared memory variables correspond to the same kernel function. The global shared memory variables are nodes in the interference graph, and the connecting edges in the interference graph characterize that the kernel function groups corresponding to the two connected global shared memory variables respectively include at least one identical kernel function.
[0050] In the embodiment of the present disclosure, each global shared memory variable can be used as a node in an interference graph. Then, based on the same kernel function in the kernel function group corresponding to every two global shared memory variables, the nodes representing the two global shared memory variables are connected in the interference graph with the kernel function as an edge, thereby generating an interference graph corresponding to the global shared memory variables.
[0051] Below Figure 2 Take the generation of interference pattern as an example to illustrate. Figure 2 A schematic diagram of an interference graph provided in an embodiment of the present disclosure.
[0052] The kernel function group corresponding to the global shared memory variable A includes kernel0 and kernel1, and the kernel function group corresponding to the global shared memory variable C includes kernel0, kernel1 and kernel2. In the interference graph, the global shared memory variables A and C can be used as a node respectively. Since the kernel function groups corresponding to the two global shared memory variables contain two identical kernel functions kernel0 and kernel1, two edges can be connected between the two nodes corresponding to the global shared memory variables A and C, which are kernel0 and kernel1 respectively, so that an interference graph can be generated, such as Figure 2 shown.
[0053] It should be noted that in Figure 2 In this paper, we only use a schematic illustration of how to generate an interference graph containing two nodes, and cannot Figure 2 As a limitation of the present disclosure, the interference graph may contain a greater number of nodes, and the manner of determining the connecting edges between these nodes is the same as described above.
[0054] It is understandable that there may be one or more connecting edges between each two nodes in the interference graph, indicating that the global shared memory variables corresponding to the two nodes are called by one or more identical kernel functions. Alternatively, there may be no connecting edges, indicating that the global shared memory variables corresponding to the two nodes are not called by the same kernel function.
[0055] Step 103: group the global shared memory variables based on the interference graph, and determine the starting position and alignment information corresponding to the global shared memory variables included in each group.
[0056] It should be noted that when two global shared memory variables are called by the same kernel function, it means that the use of the two global shared memories interferes with each other and the same address cannot be allocated. Therefore, when allocating addresses by group, the two global shared memory variables should be allocated to different groups.
[0057] In the disclosed embodiment, whether any two global shared memory variables are in a mutual interference relationship can be determined based on the connection status between the nodes in the interference graph. The global shared memory variables that interfere with each other are divided into different groups, and the global shared memory variables that do not interfere with each other are divided into the same group. In this way, all global shared memory variables can be grouped, so that the address allocation of the same group can be reused, saving shared memory space overall.
[0058] It should be noted that, in the embodiment of the present disclosure, a graph coloring algorithm may be used to color and group the nodes in the interference graph, and the global shared memory variables corresponding to the nodes with the same coloring may be determined as a group.
[0059] In the disclosed embodiment, in order to improve the memory access speed, variables in the same group need to be aligned according to boundaries when assigning addresses. Therefore, the maximum value of the alignment information corresponding to all global shared memory variables in each group can be determined as the alignment information of the group based on the alignment information contained in the declaration of each shared memory variable, so that after assigning addresses to the variables of the group, each variable can meet the boundary alignment.
[0060] In the disclosed embodiment, after determining the alignment information of each group, the starting position of each group can be allocated in descending order of the alignment information. Afterwards, within each group, each global shared memory variable can be allocated the starting position and alignment information of the group to ensure that the global shared memory variables within the same group reuse the allocated addresses.
[0061] Step 104 , determining the end address of the local shared memory variable according to the total size of all global shared memory variables and the size of the local shared memory variable.
[0062] The total size of all global shared memory variables is the sum of the maximum memory occupied by the global shared memory variables in each group after the global shared memory variables in each group are aligned.
[0063] In the disclosed embodiment, previously allocated global shared memory variables and local shared memory variables can be combined. In each kernel function, address allocation is performed in the order of global shared memory variables first and local shared memory variables later. That is, global shared memory variables are allocated starting from address 0. The starting address of the local shared memory variable is the ending address of all global shared memory variables. The size of the local shared memory variable plus the total size of the aligned global shared memory variables can obtain the final address of the local shared memory variable.
[0064] It should be noted that, when a kernel function does not call any global shared memory variables, the addresses of local shared memory variables can be allocated starting from address 0.
[0065] It should be noted that in the embodiment of the present application, local shared memory variables need to be assigned addresses only when they are used by kernel functions, while global shared memory variables need to be assigned the same address in each kernel function. This is because global shared memory variables are regarded as a whole, and each kernel function must keep an identical copy to ensure that the addresses of these global shared memory variables are consistent.
[0066] For example, when shared memory variables A and C are called by kernel functions kernel0 and kernel1 at the same time, and shared memory variable B is only called by kernel1, that is, shared memory variables A and C are global shared memory variables (global sharedvar), and shared memory variable B is a local shared memory variable (local shared var), the address allocation of shared memory variables A, B and C can be as follows: Figure 3 As shown, Figure 3 It is a schematic diagram of the address allocation result.
[0067] exist Figure 3 In the example, intA, intB, and intC represent shared memory variables. Since shared memory variables are integer variables, according to the alignment information, the starting position of each shared memory variable should be an integer multiple of 4. Both intA and intC are global shared memory variables and interfere with each other. They can be determined as a group and assigned an address. According to the allocation order of global shared memory variables first and local shared memory variables later, it can be determined that the starting position of intA is 0 and the starting position of intC is 4, and the addresses allocated in kernel functions kernel0 and kernel1 are consistent. Since intB is only called by kernel1, it only needs to allocate an address in the shared memory of kernel1. The starting position of intB is 8 and the ending address is 12.
[0068] In the disclosed embodiment, all shared memory variables are first classified into local shared memory variables and global shared memory variables, and then the global shared memory variables are grouped according to the kernel function group corresponding to the global shared memory variables, and the starting position and alignment information of each group in the memory space are determined to complete the address allocation of the global shared memory variables. Afterwards, according to the total size of the global shared memory variables after group alignment and the size of the local shared memory variables, the end address of the local shared memory variables is determined to complete the address allocation of the local shared memory variables. Thereby, the address allocation of the global shared memory variables is optimized, the redundancy of the address allocation is effectively reduced, and the memory space consumption is reduced.
[0069] The following is an exemplary description of the shared memory variable classification, global shared memory variable grouping and address allocation proposed in the embodiment of the present disclosure in conjunction with FIG. 4 .
[0070] First, the shared memory variables include a, b, c, d, and e. The call information of each shared memory variable is as follows: Figure 4a As shown, Figure 4a It is a diagram of how a kernel function calls a device-side function and how the device-side function references a shared memory variable. Figure 4aIt can be seen that the shared memory variable a is declared in the device-side function func1, the shared memory variable b is declared in the device-side function func2, the shared memory variable c is declared in the device-side function func3, the shared memory variable d is declared in the device-side function func4, and the shared memory variable e is declared in the device-side function func5. In addition, the kernel function kernel1 calls func1, the kernel function kernel2 calls func1, func4 and func5, the kernel functions kernel3 and kernel4 both call func1, func2 and func3, and the kernel function kernel5 calls func2 and func4.
[0071] Since the device-side functions that reference shared memory variables a, b, c, d, and e are all called by two or more kernel functions, it can be determined that shared memory variables a, b, c, d, and e are all global shared memory variables. Figure 4a It can be seen that func3 and func4 are not used on the same kernel, and the global shared memory variables c and d do not interfere with each other in address allocation, so c and d can be grouped together. Similarly, func3 and func5 are not used on the same kernel, and func2 and func5 are not used on the same kernel. c and e or b and e can also be grouped together.
[0072] It should be noted that c and d can be a group, c and e can also be a group, but d and e cannot be a group. In the present disclosure, it can be determined according to the order of determining the grouping. That is, when it is determined that c and d can be a group, c and d are directly determined as a group. If it is determined that c and e can be a group later, it is determined whether e and d can be a group. If not, e is placed in another group.
[0073] Since the device-side function func1 corresponding to the global shared memory variable a has the same calling kernel function as any other device-side function, the global shared memory variable a can be identified as a group separately. The grouping of the global shared memory variables is as follows: Figure 4b As shown, the global shared memory variables a, b, c, d and e can be divided into three groups. The first group group1 contains the global shared memory variable a, the second group group2 contains the global shared memory variables b and e, and the third group group3 contains the global shared memory variables c and d.
[0074] After that, the size and alignment information of the global shared memory variables in each group can be taken to their maximum values to obtain the size corresponding to each group (i.e. Figure 4b The value corresponding to size in) and alignment information (i.e. Figure 4bIn the example above, if the variables in each group are all integers (int), that is, the size is 4 bytes, and the alignment information is aligned to a 4-byte boundary, it can be determined that the size and alignment information corresponding to each group are both 4. Figure 4b As shown, the starting position of each group can be determined in turn to complete the address allocation of the global shared memory variable. The allocation result is as follows Figure 4c shown.
[0075] It should be noted that Figure 4c Schematic diagram of the allocation result corresponding to the address allocation method proposed in this disclosure. Figure 4d Compared with the schematic diagram of the allocation result obtained after address allocation for shared memory variables in the prior art, the address allocation scheme disclosed in the present invention can make shared memory variables share address allocation, and from the final address arrangement, the usage of shared memory in each kernel is reduced by 8 bytes. Therefore, it is shown that the address allocation method proposed in the present invention can effectively save memory space and improve the performance of address allocation.
[0076] Figure 5 A flow chart of a method for allocating addresses of shared memory variables provided by an embodiment of the present disclosure is shown as follows: Figure 5 As shown, the address allocation method of the shared memory variable may include the following steps:
[0077] Step 501 : Determine, according to the calling information of each shared memory variable, the local shared memory variable, the global shared memory variable and the kernel function group corresponding to each global shared memory variable.
[0078] In the embodiment of the present disclosure, the shared memory variable can be directly called by the kernel function, or the kernel function can also call the shared memory variable by calling the device-side function that references the shared memory variable. Therefore, when the shared memory variables are classified into two categories according to the call information of each shared memory variable, at least one of the following items may be included:
[0079] Optionally, the calling information of each shared memory variable may be traversed, and when any shared memory variable is called by only one kernel function, the shared memory variable is determined to be a local shared memory variable;
[0080] Alternatively, the calling information of each shared memory variable may be traversed, and when any shared memory variable is not called by any kernel function and the function where the shared memory variable is located is only called by one kernel function, the shared memory variable is determined to be a local shared memory variable.
[0081] Combine the following Figure 5 Explain the classification of shared memory variables. Figure 5A schematic diagram of a shared memory variable classification process provided by the present disclosure.
[0082] Depend on Figure 5 It can be seen that, for the shared memory variables to be classified, it can be determined whether the reference (or call) of the variable is only in a single kernel function. If it is determined that it is only called in a single kernel function, the variable is determined as a local shared memory variable. Otherwise, it is further determined whether the function where the variable references is called by a single kernel function. If it is only called by a single kernel function, the variable can be determined as a local shared memory variable. Otherwise, the variable is determined as a global shared memory variable.
[0083] In the disclosed embodiment, the classification of shared memory variables is determined by judging the shared memory variables and the calling situations of the functions where the shared memory variables are located by the kernel functions, which can ensure the accuracy and reliability of the shared memory variable classification results and provide a reliable data basis for subsequent address allocation.
[0084] Step 502: Generate an interference graph corresponding to each global shared memory variable based on the kernel function group corresponding to each global shared memory variable.
[0085] For detailed description of the above steps 501 and 502, please refer to other embodiments of the present disclosure and will not be repeated here.
[0086] Step 503: determine the global shared memory variables without connecting edges in the interference graph as a group.
[0087] In the embodiment of the present disclosure, in the interference graph, global shared memory variables without connecting edges can be used to determine that the use of these global shared memories will not interfere with each other, and the same address can be assigned. Therefore, these global shared memory variables can be determined as a group to reuse the same address.
[0088] It should be noted that, because the interference graph is generated based on global shared memory variables, each global shared memory variable has at least one connection edge in the interference graph. When any global shared memory variable has a connection edge with all other global shared memory variables in the interference graph, the any global shared memory variable can be determined as a group. In the present disclosure, all global shared memory variables can be divided into one or more groups.
[0089] For example, the global shared memory variables are A, B, C, D and E. In the interference graph, A is connected to B, C, D and E, B and E do not have a connecting edge, and C and D do not have a connecting edge. Then A can be determined as a group, B and E can be determined as a group, and C and D can be determined as a group.
[0090] Step 504 : Determine alignment information corresponding to each group according to alignment information corresponding to each global shared memory variable included in each group.
[0091] In the disclosed embodiment, the heterogeneous device program can determine the alignment information corresponding to each global shared memory variable contained in each group based on the alignment information contained in each shared memory variable declaration, and then the alignment information corresponding to all the global shared memory variables contained in the group can be sorted to determine the alignment information corresponding to the group.
[0092] Optionally, the maximum value of the alignment information corresponding to all global shared memory variables contained in each group can be determined as the alignment information corresponding to the group. This ensures that the alignment information corresponding to each group can satisfy the alignment information corresponding to all global shared memory variables in the group, thereby ensuring the correct alignment and efficient access of data in the memory, as well as the integrity of the data, and optimizing the memory layout.
[0093] Step 505: Determine the starting position corresponding to each group according to the alignment information corresponding to each group.
[0094] In the embodiment of the present disclosure, the starting position corresponding to each group may be determined in sequence according to the order of alignment information corresponding to each group from large to small.
[0095] Optionally, the starting position of the shared memory may be first determined as the starting position corresponding to a group having the largest value in the corresponding alignment information.
[0096] In the disclosed embodiment, in order to optimize memory layout and access efficiency and reduce the complexity of memory allocation and management, the starting position corresponding to each group can be allocated in descending order of alignment information. Therefore, the starting position of the shared memory (marked as 0) can be first determined as the starting position corresponding to the group with the largest value in the corresponding alignment information.
[0097] Then, the remaining groups may be sorted in descending order according to the alignment information, and the starting position corresponding to each group in the remaining groups may be determined in turn.
[0098] In the disclosed embodiment, the remaining groups can be sorted in descending order of the alignment information according to the rules that the starting position of each group in the memory must meet (such as an integer multiple of a fixed value (such as 2, 4, 8, etc.) specified by the alignment information of each group, and the starting position corresponding to each group in the remaining groups can be determined in turn. This can improve the orderliness of address allocation, make the address allocation result more reliable, and avoid memory allocation failure.
[0099] Step 506: determine the alignment information and the starting position corresponding to each group as the alignment information and the starting position corresponding to the global shared memory variable included in the group.
[0100] In the disclosed embodiment, since all global shared memory variables in the same group reuse the same address, the alignment information and starting position corresponding to each group can be determined as the alignment information and starting position corresponding to the global shared memory variables included in the group.
[0101] Step 507 , determining the end address of the local shared memory variable according to the total size of all global shared memory variables and the size of the local shared memory variable.
[0102] For detailed description of the above step 507, please refer to other embodiments of the present disclosure and will not be repeated here.
[0103] In this embodiment, the alignment information corresponding to each group is determined based on the alignment information corresponding to the global shared memory variables, and then the starting position of each group is determined in turn according to the size of the alignment information, and the alignment information and starting position of the group are used to unify the alignment information and starting positions corresponding to all shared memory variables in the group. Thus, the reliability and accuracy of the address allocation result can be improved, ensuring that all global shared memory variables in the same group can reuse the same address, thereby improving the utilization rate of the memory space.
[0104] Figure 7 A flow chart of a method for allocating addresses of shared memory variables provided by an embodiment of the present disclosure is shown as follows: Figure 7 As shown, the address allocation method of the shared memory variable may include the following steps:
[0105] Step 701: Determine the local shared memory variable, the global shared memory variable, and the kernel function group corresponding to each global shared memory variable according to the calling information of each shared memory variable.
[0106] Step 702: Generate an interference graph corresponding to each global shared memory variable based on the kernel function group corresponding to each global shared memory variable.
[0107] Step 703: group the global shared memory variables based on the interference graph, and determine the starting position and alignment information corresponding to the global shared memory variables included in each group.
[0108] For detailed description of the above steps 701 to 703, reference can be made to other embodiments of the present disclosure and will not be repeated here.
[0109] Step 704: when there are multiple local shared memory variables, obtain alignment information, sizes, and initial appearance order of the multiple local shared memory variables.
[0110] In the disclosed embodiment, for local shared memory variables, addresses of these variables can be respectively allocated in the kernel function that calls these local shared memory variables. Therefore, in the case of multiple local shared memory variables, the address allocation order of these local shared memory variables needs to be considered to ensure optimal address allocation and save memory space as much as possible, so the alignment information, size and initial appearance order of multiple local shared memory variables can be obtained.
[0111] Step 705 , sorting the multiple local shared memory variables based on the alignment information, size and initial appearance order of the multiple local shared memory variables, and determining the starting position of each local shared memory variable.
[0112] In the disclosed embodiment, a structure element optimization sorting algorithm may be used to sort multiple local shared memory variables in combination with the alignment information, size, and initial appearance order priority of the multiple local shared memory variables.
[0113] Optionally, multiple local shared memory variables can be sorted in descending order of alignment information. Then, when the alignment information corresponding to any two local shared memory variables is the same, the two local shared memory variables can be sorted in descending order of size. Afterwards, when the sizes corresponding to any two local shared memory variables are the same, the two local shared memory variables can be sorted in the initial order of appearance until the starting positions corresponding to all local shared memory variables are determined. Thereby, the orderliness of the address allocation of the local shared memory variables can be improved, the effect of address allocation can be optimized, and it is beneficial to save memory space.
[0114] It should be noted that after determining the starting positions corresponding to all local shared memory variables, it is possible to detect whether there are gaps in the addresses allocated at this time. Because gaps occupy memory space, these spaces are not effectively utilized, which will lead to waste of memory resources and reduce access efficiency. Therefore, when there are no gaps, the address allocation of local shared memory variables can be terminated. Otherwise, when there are gaps, it is necessary to use methods such as greedy algorithms to reduce the gaps.
[0115] Optionally, in the case where there are gaps between multiple local shared memory variables, the local shared memory variable with the largest alignment information and the largest size may be first determined as the first local shared memory variable.
[0116] Then, based on the alignment information and sizes of the remaining local shared memory variables, the remaining local shared memory variables may be traversed until a second local shared memory variable having no gap with the first local shared memory variable is determined.
[0117] In the disclosed embodiment, traversal can be performed in reverse order of the alignment information and sizes of the remaining local shared memory variables, and it is determined in turn whether there is a gap between each local shared memory variable and the first local shared memory variable determined above. When it is determined that there is no gap between a local shared memory variable and the first local shared memory variable, the local shared memory variable can be directly determined as the second local shared memory variable, and the traversal is terminated.
[0118] It should be noted that after traversing the remaining local shared memory variables, there may be a gap between each remaining local shared memory variable and the first local shared memory variable, in which case the local shared memory variable with the smallest gap may be determined as the second local shared memory variable. Alternatively, if the gaps between all remaining local shared memory variables and the first local shared memory variable are the same, the remaining shared memory variable with the largest alignment information and size may be taken as the second local shared memory variable.
[0119] Afterwards, the operation of traversing the remaining local shared memory variables based on the alignment information and sizes of the remaining local shared memory variables is returned until the order of all local shared memory variables is determined.
[0120] It can be understood that each time the next local shared memory variable is determined, the traversal is completed by calculating the gap with the last local shared memory variable determined previously.
[0121] Therefore, after determining the order of all local shared memory variables, the starting position corresponding to each local shared memory variable can be re-determined according to the order of all local shared memory variables and the alignment information of each local shared memory variable, so that the addresses allocated to multiple local shared memory variables are as close as possible, the analysis between addresses is reduced, the effect of address allocation is further optimized, and memory space is saved.
[0122] Step 706, determining the end address of the local shared memory variable according to the total size of all global shared memory variables and the size of the local shared memory variable.
[0123] For a detailed description of the above step 706, please refer to other embodiments of the present disclosure and will not be repeated here.
[0124] In this embodiment, multiple local shared memory variables are sorted based on the alignment information, size and initial appearance order of the local shared memory variables, and addresses are allocated in sequence, thereby reducing the gaps between the addresses allocated to the local shared memory variables, saving memory space, and optimizing the effect of address allocation.
[0125] In order to implement the above embodiment, the present disclosure also proposes an address allocation device for shared memory variables.
[0126] Figure 8 A schematic diagram of the structure of a device for allocating addresses for shared memory variables provided in an embodiment of the present disclosure.
[0127] like Figure 8 As shown, the shared memory variable address allocation device 800 may include:
[0128] A first determination module 801 is used to determine a local shared memory variable, a global shared memory variable and a kernel function group corresponding to each global shared memory variable according to the call information of each shared memory variable;
[0129] A generating module 802 is used to generate an interference graph corresponding to each global shared memory variable based on the kernel function group corresponding to each global shared memory variable, wherein the global shared memory variable is a node in the interference graph, and the connecting edge in the interference graph indicates that the kernel function groups corresponding to two connected global shared memory variables respectively include at least one identical kernel function;
[0130] A second determination module 803 is used to group the global shared memory variables based on the interference graph, and determine the starting position and alignment information corresponding to the global shared memory variables included in each group;
[0131] The third determining module 804 is used to determine the end address of the local shared memory variable according to the total size of all global shared memory variables and the size of the local shared memory variable.
[0132] In some possible embodiments, the first determining module 801 may be specifically used for at least one of the following:
[0133] Traversing the call information of each shared memory variable, if any shared memory variable is called by only one kernel function, determining that any shared memory variable is a local shared memory variable;
[0134] The calling information of each shared memory variable is traversed, and when any shared memory variable is not called by any kernel function and the function where the any shared memory variable is located is only called by one kernel function, it is determined that the any shared memory variable is a local shared memory variable.
[0135] In some possible embodiments, the second determining module 803 may be specifically configured to:
[0136] The global shared memory variables that do not have connected edges in the interference graph are identified as a group;
[0137] Determine alignment information corresponding to the group according to alignment information corresponding to each global shared memory variable included in each group;
[0138] According to the alignment information corresponding to each group, determine the starting position corresponding to each group;
[0139] The alignment information and the starting position corresponding to each group are determined as the alignment information and the starting position corresponding to the global shared memory variables contained in the group.
[0140] In some possible embodiments, the second determining module 803 may be specifically configured to:
[0141] The maximum value of the alignment information corresponding to all the global shared memory variables contained in each group is determined as the alignment information corresponding to the group.
[0142] In some possible embodiments, the second determining module 803 may be specifically configured to:
[0143] The starting position of the shared memory is determined as the starting position corresponding to the group with the largest value in the corresponding alignment information;
[0144] The remaining groups are sorted in descending order according to the alignment information, and the starting position corresponding to each group in the remaining groups is determined in turn.
[0145] In some possible embodiments, the third determining module 804 may also be used to:
[0146] In the case where there are multiple local shared memory variables, the alignment information, size and initial appearance order of the multiple local shared memory variables are obtained;
[0147] Based on the alignment information, size and initial appearance order of the multiple local shared memory variables, the multiple local shared memory variables are sorted, and the starting position of each local shared memory variable is determined respectively.
[0148] In some possible embodiments, the third determining module 804 may be specifically configured to:
[0149] Sort multiple local shared memory variables in descending order according to the alignment information;
[0150] When the alignment information corresponding to any two local shared memory variables is the same, the two local shared memory variables are sorted in descending order of size;
[0151] When any two local shared memory variables have the same size, the two local shared memory variables are sorted according to the initial appearance order until the starting positions corresponding to all local shared memory variables are determined.
[0152] In some possible embodiments, the third determining module 804 may also be used to:
[0153] In the case where there are gaps between multiple local shared memory variables, the local shared memory variable with the largest alignment information and the largest size is determined as the first local shared memory variable;
[0154] Based on the alignment information and sizes of the remaining local shared memory variables, traverse the remaining local shared memory variables until a second local shared memory variable having no gap with the first local shared memory variable is determined;
[0155] Return to perform the operation of traversing the remaining local shared memory variables based on the alignment information and size of the remaining local shared memory variables until the order of all local shared memory variables is determined;
[0156] According to the order of all local shared memory variables and the alignment information of each local shared memory variable, the starting position corresponding to each local shared memory variable is re-determined.
[0157] The functions and specific implementation principles of the above modules in the embodiments of the present disclosure can be referred to the above method embodiments, and will not be repeated here.
[0158] The address allocation device for shared memory variables of the disclosed embodiment first classifies all shared memory variables into local shared memory variables and global shared memory variables, then groups the global shared memory variables according to the kernel function groups corresponding to the global shared memory variables, determines the starting position and alignment information of each group in the memory space, and completes the address allocation for the global shared memory variables. Afterwards, according to the total size of the global shared memory variables after group alignment and the size of the local shared memory variables, the end address of the local shared memory variables is determined to complete the address allocation for the local shared memory variables. Thereby, the address allocation for the global shared memory variables is optimized, the redundancy of the address allocation is effectively reduced, and the memory space consumption is reduced.
[0159] In order to implement the above embodiments, the present disclosure further proposes an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the address allocation method for shared memory variables proposed in the above embodiments of the present disclosure is implemented.
[0160] Fig. 9 A block diagram of an exemplary electronic device suitable for implementing embodiments of the present disclosure is shown. Fig. 9 The electronic device 12 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0161] like Fig. 9As shown, the electronic device 12 is in the form of a general purpose computing device. The components of the electronic device 12 may include, but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 that connects various system components (including the system memory 28 and the processing unit 16).
[0162] The bus 18 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor or a local bus using any of a variety of bus structures. For example, these architectures include but are not limited to Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus and Peripheral Component Interconnection (PCI) bus.
[0163] The electronic device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the electronic device 12, including volatile and non-volatile media, removable and non-removable media.
[0164] The memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The electronic device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 34 may be used to read and write non-removable, non-volatile magnetic media ( Fig. 9 not shown, usually called a "hard drive"). Although Fig. 9Not shown in the figure, a disk drive for reading and writing a removable non-volatile disk (e.g., a "floppy disk"), and an optical disk drive for reading and writing a removable non-volatile optical disk (e.g., a compact disc read only memory (CD-ROM), a digital versatile disc read only memory (DVD-ROM), or other optical media) may be provided. In these cases, each drive may be connected to the bus 18 via one or more data medium interfaces. The memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the various embodiments of the present disclosure.
[0165] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in the memory 28, such program modules 42 including but not limited to an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment. The program modules 42 generally perform the functions and / or methods of the embodiments described in the present disclosure.
[0166] The electronic device 12 may also communicate with one or more external devices 14 (e.g., keyboards, pointing devices, displays 24, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 12, and / or may communicate with any device that enables the electronic device 12 to communicate with one or more other computing devices (e.g., network cards, modems, etc.). Such communication may be performed via an input / output (I / O) interface 22. Furthermore, the electronic device 12 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 20. As shown, the network adapter 20 communicates with other modules of the electronic device 12 via a bus 18. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0167] The processing unit 16 executes various functional applications and data processing by running programs stored in the system memory 28, such as implementing the methods mentioned in the above embodiments.
[0168] In order to implement the above embodiments, the present disclosure further proposes a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the address allocation method of the shared memory variable proposed in the above embodiments of the present disclosure is implemented.
[0169] In order to implement the above embodiments, the present disclosure also proposes a chip, which includes a processing circuit and an interface circuit; wherein the interface circuit is used to obtain instructions and send the instructions to the processing circuit, and the processing circuit is used to execute instructions to implement the address allocation method of shared memory variables proposed in the above embodiments of the present disclosure.
[0170] Fig.10 is a schematic diagram of the structure of the chip proposed in the embodiment of the present disclosure. Fig.10 The structure of the chip 1000 is shown, but is not limited to this.
[0171] The chip 1000 includes a processing circuit 1001 , and the processing circuit 1001 is configured to execute any of the above methods.
[0172] In some embodiments, the chip 1000 further includes one or more interface circuits 1002. Optionally, the interface circuit 1002 is connected to the memory 1003. The interface circuit 1002 can be used to receive signals from the memory 1003 or other devices, and the interface circuit 1002 can be used to send signals to the memory 1003 or other devices. For example, the interface circuit 1002 can read instructions stored in the memory 1003 and send the instructions to the processing circuit 1001.
[0173] In some embodiments, the interface circuit 1002 performs at least one of the communication steps such as sending and / or receiving in the above method, and the processing circuit 1001 performs other steps.
[0174] In some embodiments, terms such as interface circuit, interface, transceiver pin, and transceiver may be used interchangeably.
[0175] In some embodiments, the chip 1000 further includes one or more memories 1003 for storing instructions. Optionally, all or part of the memory 1003 may be outside the chip 1000.
[0176] The technical solution disclosed in the present invention can classify and group shared memory variables according to the calling conditions of the shared memory variables in different kernel functions, and then allocate addresses, so that some shared memory variables can reuse the same address, optimize the space allocation of shared memory variables, and reduce space consumption.
[0177] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0178] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of the present disclosure, "plurality" means at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0179] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code that includes one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present disclosure includes additional implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present disclosure belong.
[0180] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute the instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purpose of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing in other suitable ways if necessary, and then stored in a computer memory.
[0181] It should be understood that the various parts of the present disclosure can be implemented in hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented in software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0182] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.
[0183] In addition, each functional unit in each embodiment of the present disclosure may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0184] The storage medium mentioned above may be a read-only memory, a magnetic disk or an optical disk, etc. Although the embodiments of the present disclosure have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations of the present disclosure. A person of ordinary skill in the art may change, modify, replace and modify the above embodiments within the scope of the present disclosure.
Claims
1. A method for allocating addresses of shared memory variables, characterized in that: include: Determine the local shared memory variable, the global shared memory variable and the kernel function group corresponding to each global shared memory variable according to the calling information of each shared memory variable; Based on the kernel function group corresponding to each global shared memory variable, an interference graph corresponding to the global shared memory variable is generated, wherein the global shared memory variable is a node in the interference graph, and the connecting edge in the interference graph represents that the kernel function groups corresponding to two connected global shared memory variables respectively include at least one identical kernel function; Based on the interference graph, the global shared memory variables are grouped, and the starting positions and alignment information corresponding to the global shared memory variables included in each group are determined; According to the total size of all global shared memory variables and the size of the local shared memory variable, the termination address of the local shared memory variable is determined.
2. The method according to claim 1, characterized in that The determining, according to the calling information of each shared memory variable, the local shared memory variable, the global shared memory variable and the kernel function group corresponding to each global shared memory variable comprises at least one of the following: Traversing the call information of each shared memory variable, and determining that any shared memory variable is a local shared memory variable when any shared memory variable is called by only one kernel function; The calling information of each shared memory variable is traversed, and when any shared memory variable is not called by any kernel function and the function where the any shared memory variable is located is only called by one kernel function, it is determined that the any shared memory variable is a local shared memory variable.
3. The method according to claim 1, characterized in that The method of grouping the global shared memory variables based on the interference graph and determining the starting position and alignment information corresponding to the global shared memory variables included in each group includes: Determine the global shared memory variables that do not have connected edges in the interference graph as a group; Determine alignment information corresponding to the group according to alignment information corresponding to each global shared memory variable included in each group; According to the alignment information corresponding to each group, determine the starting position corresponding to each group; The alignment information and the starting position corresponding to each group are determined as the alignment information and the starting position corresponding to the global shared memory variables contained in the group.
4. The method according to claim 3, characterized in that The step of determining the alignment information corresponding to each group according to the alignment information corresponding to each global shared memory variable included in each group includes: The maximum value of the alignment information corresponding to all the global shared memory variables contained in each group is determined as the alignment information corresponding to the group.
5. The method according to claim 3, characterized in that The step of determining the starting position corresponding to each group according to the alignment information corresponding to each group includes: The starting position of the shared memory is determined as the starting position corresponding to the group with the largest value in the corresponding alignment information; The remaining groups are sorted in descending order according to the alignment information, and the starting position corresponding to each group in the remaining groups is determined in turn.
6. The method according to any one of claims 1 to 5, characterized in that: Before determining the termination address of the local shared memory variable according to the total size of all global shared memory variables and the size of the local shared memory variable, the method further includes: In the case where there are multiple local shared memory variables, obtaining alignment information, sizes, and initial appearance order of the multiple local shared memory variables; Based on the alignment information, size and initial appearance order of the multiple local shared memory variables, the multiple local shared memory variables are sorted, and the starting position of each of the local shared memory variables is determined respectively.
7. The method according to claim 6, characterized in that The step of sorting the multiple local shared memory variables based on the alignment information, size and initial appearance order of the multiple local shared memory variables and determining the starting position of each of the local shared memory variables comprises: Sort the multiple local shared memory variables in descending order of alignment information; When the alignment information corresponding to any two local shared memory variables are the same, the two local shared memory variables are sorted in descending order of size; In the case that the sizes corresponding to any two local shared memory variables are the same, the any two local shared memory variables are sorted according to the initial appearance order until the starting positions corresponding to all local shared memory variables are determined.
8. The method according to claim 7, characterized in that After determining the starting positions corresponding to all local shared memory variables respectively, the method further includes: In the case where there are gaps between the multiple local shared memory variables, determining the local shared memory variable with the largest alignment information and the largest size as the first local shared memory variable; Based on the alignment information and the size of the remaining local shared memory variables, traverse the remaining local shared memory variables until a second local shared memory variable having no gap with the first local shared memory variable is determined; Returning to execute the operation of traversing the remaining local shared memory variables based on the alignment information and the size of the remaining local shared memory variables until the order of all local shared memory variables is determined; According to the order of all the local shared memory variables and the alignment information of each of the local shared memory variables, a starting position corresponding to each of the local shared memory variables is re-determined.
9. A device for allocating addresses of shared memory variables, characterized in that: The device comprises: A first determination module is used to determine the local shared memory variable, the global shared memory variable and the kernel function group corresponding to each global shared memory variable according to the call information of each shared memory variable; A generating module, configured to generate an interference graph corresponding to the global shared memory variable based on the kernel function group corresponding to each global shared memory variable, wherein the global shared memory variable is a node in the interference graph, and the connecting edge in the interference graph indicates that the kernel function groups corresponding to two connected global shared memory variables respectively include at least one identical kernel function; A second determination module is used to group the global shared memory variables based on the interference graph, and determine the starting position and alignment information corresponding to the global shared memory variables included in each group; The third determining module is used to determine the end address of the local shared memory variable according to the total size of all global shared memory variables and the size of the local shared memory variable.
10. An electronic device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the program, the method for allocating addresses of shared memory variables as claimed in any one of claims 1 to 8 is implemented.
11. A chip, characterized in that: The chip includes a processing circuit and an interface circuit; wherein the interface circuit is used to obtain instructions and send the instructions to the processing circuit, and the processing circuit is used to execute the instructions to implement the address allocation method for shared memory variables as described in any one of claims 1-8.
12. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for allocating addresses of shared memory variables as described in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Memory sharing method and related device
CN116560878A
Sample component analysis method and system, storage medium and computer
CN117473233A
Hardware Memory Error Tolerant Software System
US20230185663A1