Shared memory access method and device, equipment, storage medium and program product

By judging and splitting the same specification information in the shared memory specification instruction in the parallel calculation of the general graphics processor, the problem of multiple read and write channels with the same address in the workgroup thread Wave is solved, and the performance of shared memory specification operation is improved.

CN119961019APending Publication Date: 2025-05-09GLENFLY TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510025176.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

In parallel computing of general graphics processors, channels with the same address of the Reduction instruction in the workgroup thread Wave need to perform multiple read operations and write operations, resulting in reduced performance of shared memory specification operations.

Method used

By receiving shared memory specification instructions, the same specification judgment is made on the valid channel based on the address information and splitting it, reducing the number of read and write times to the shared memory. The specific steps include receiving regulations and determining whether the channel address is the same. If the same is true, merge. If the XOR operation result is 0, the determination address is the same. If the XOR result is not 0, the determination address is different. If the channel is split according to the result, the read and write operations will be reduced.

Benefits of technology

By reducing the number of read and write times of shared memory, the efficiency of shared memory data specifications is significantly improved, suitable for shared memory in different configurations, different protocol opcodes, and support for different data formats.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961019A_ABST
    Figure CN119961019A_ABST
Patent Text Reader

Abstract

The invention relates to a shared memory access method and device, equipment, a storage medium and a program product. The method comprises the following steps: receiving a shared memory protocol instruction; according to address information contained in the shared memory protocol instruction, same protocol judgment is conducted on the effective channels, and the effective channels are split according to the same protocol judgment result; reading data from the shared memory according to the splitting result of the effective channel; wherein the splitting result is associated with the number of times of reading the data from the shared memory; and performing protocol operation according to the data read in the shared memory, and the source data and the operation code contained in the shared memory protocol instruction, and writing a protocol operation result back to the shared memory. Therefore, the effective channels with the same address can be merged, the read-write frequency of the shared memory is reduced, the efficiency of the shared memory data protocol is remarkably improved, and the method and the device can be suitable for shared memories with different configurations and different protocol operation codes and support different data formats.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method, apparatus, computer device, computer-readable storage medium and computer program product for accessing a shared memory. Background Art

[0002] With the rapid development of computer technology, a large number of parallel processing units contained in graphics processors can be programmed to perform parallel computing of single instruction multiple data streams, for example: using general-purpose image processors for parallel computing.

[0003] In traditional technology, in parallel computing of general-purpose graphics processors, work items of the same work group can access the same shared memory space and perform reduction operations or atomic operations. Among them, reduction operations need to be performed in sequence, and the address of each channel in the work group thread Wave (a group of thread blocks that can be executed concurrently) must first perform a shared memory read, then perform a reduction operation, and finally write the result to the shared memory through a shared memory write operation.

[0004] However, for the lanes with the same address as the Reduction instructions in the workgroup thread Wave, multiple read and write operations are required, resulting in a decrease in the performance of the shared memory reduction operation. Summary of the invention

[0005] Based on this, it is necessary to provide a shared memory access method, device, computer equipment, computer-readable storage medium and computer program product that can reduce the number of read and write times of protocol operations on shared memory and improve the performance of shared memory protocol operations in response to the above technical problems.

[0006] In a first aspect, the present application provides a method for accessing a shared memory, comprising:

[0007] Receive a shared memory protocol instruction; the instruction information contained in the shared memory protocol instruction includes: address information, operation code, source data;

[0008] According to the address information contained in the shared memory protocol instruction, the valid channel is judged to have the same protocol, and the valid channel is split according to the result of the same protocol judgment; the address information is used to indicate the address of the valid channel;

[0009] Reading data from the shared memory according to the splitting result of the valid channel; wherein the splitting result is associated with the number of times the data needs to be read from the shared memory;

[0010] A reduction operation is performed according to the data read from the shared memory, the source data and the operation code contained in the shared memory reduction instruction, and the result of the reduction operation is written back to the shared memory.

[0011] In one embodiment, receiving a shared memory protocol instruction includes:

[0012] receiving a shared memory protocol instruction sent by an arithmetic logic operation unit;

[0013] The instruction information contained in the shared memory protocol instruction is stored in the shared memory buffer.

[0014] In one embodiment, performing same protocol determination on valid channels according to address information included in the shared memory protocol instruction, and splitting the valid channels according to the result of the same protocol determination, comprises:

[0015] Obtaining the instruction information from the shared memory buffer through a protocol instruction information control unit;

[0016] Traverse to obtain the addresses of valid channels, and perform pairwise XOR operations on the addresses of valid channels. If the result of the pairwise XOR operation is not 0, it is determined that the addresses of the two valid channels are different; if the result of the pairwise XOR operation is 0, it is determined that the addresses of the two valid channels are the same;

[0017] Merge valid channels with the same address and only keep valid channels with different addresses;

[0018] The valid channels with different addresses are split according to the channel addresses corresponding to the valid channels.

[0019] In one embodiment, the traversing to obtain addresses of valid channels and performing pairwise XOR operations on the addresses of valid channels includes:

[0020] Perform an XOR operation on the address of the i-th valid channel and the address of the i+1-th valid channel, where i is a natural number in [0, n-1] and n is the total number of valid channels.

[0021] In one embodiment, splitting the valid channels according to the channel addresses corresponding to the valid channels with different addresses includes:

[0022] If the addresses of all valid channels are the same, the number of splits is determined to be 1, and the addresses of all valid channels are stored in a storage library;

[0023] If the addresses of the valid channels are different, it is determined whether the addresses of the valid channels span different storage bins. If the addresses of the valid channels are all in one storage bin, the addresses of all the valid channels are stored in one storage bin. If the addresses of the valid channels are not all in one storage bin, the number of splits is determined according to the number of storage bins spanned by the addresses of the valid channels.

[0024] The configuration of the shared memory determines the number of storage banks and the width of each storage bank. Each storage bank corresponds to a split unit, and the number of split units corresponds to the number of times data needs to be read from the shared memory.

[0025] In one embodiment, reading data from a shared memory according to the splitting result of the valid channel includes:

[0026] Determine the number of times data needs to be read from the shared memory according to the split result of the valid channel, and determine the shared memory read request address corresponding to each read operation;

[0027] According to the shared memory read request address, corresponding data is read from the shared memory.

[0028] In one embodiment, performing a reduction operation according to the data read from the shared memory, the source data and the operation code included in the shared memory reduction instruction includes:

[0029] Match the data read from the shared memory with the valid channels indicated by each channel index to obtain the data corresponding to each valid channel;

[0030] According to the operation code included in the shared memory protocol instruction, a protocol operation is performed on the data corresponding to each valid channel and the source data corresponding to each valid channel in the shared memory protocol instruction.

[0031] In one embodiment, the instruction information included in the shared memory protocol instruction also includes: the number of valid channels, identification information of whether to write back the protocol result, and the data format supported by the protocol.

[0032] In one embodiment, after writing the result of the reduction operation back to the shared memory, the method further includes:

[0033] When the identification information of whether to write back the reduction result is not empty, the result of the reduction operation is written back to the general file register.

[0034] In a second aspect, the present application further provides a shared memory access device, comprising:

[0035] A receiving module, used for receiving a shared memory protocol instruction; the instruction information contained in the shared memory protocol instruction includes: address information, operation code, source data;

[0036] A judgment module, used for performing same-rule judgment on valid channels according to the address information contained in the shared memory protocol instruction, and splitting the valid channels according to the result of the same-rule judgment; the address information is used to indicate the address of the valid channel;

[0037] A reading module, used for reading data from the shared memory according to the splitting result of the valid channel; wherein the splitting result is associated with the number of times the data needs to be read from the shared memory;

[0038] The protocol module is used to perform a protocol operation according to the data read from the shared memory, the source data and the operation code contained in the shared memory protocol instruction, and write the result of the protocol operation back to the shared memory.

[0039] In a third aspect, the present application further provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0040] Receive a shared memory protocol instruction; the instruction information contained in the shared memory protocol instruction includes: address information, operation code, source data;

[0041] According to the address information contained in the shared memory protocol instruction, the valid channel is judged to have the same protocol, and the valid channel is split according to the result of the same protocol judgment; the address information is used to indicate the address of the valid channel;

[0042] Reading data from the shared memory according to the splitting result of the valid channel; wherein the splitting result is associated with the number of times the data needs to be read from the shared memory;

[0043] A reduction operation is performed according to the data read from the shared memory, the source data and the operation code contained in the shared memory reduction instruction, and the result of the reduction operation is written back to the shared memory.

[0044] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the following steps are implemented:

[0045] Receive a shared memory protocol instruction; the instruction information contained in the shared memory protocol instruction includes: address information, operation code, source data;

[0046] According to the address information contained in the shared memory protocol instruction, the valid channel is judged to have the same protocol, and the valid channel is split according to the result of the same protocol judgment; the address information is used to indicate the address of the valid channel;

[0047] Reading data from the shared memory according to the splitting result of the valid channel; wherein the splitting result is associated with the number of times the data needs to be read from the shared memory;

[0048] A reduction operation is performed according to the data read from the shared memory, the source data and the operation code contained in the shared memory reduction instruction, and the result of the reduction operation is written back to the shared memory.

[0049] In a fifth aspect, the present application further provides a computer program product, including a computer program, which implements the following steps when executed by a processor:

[0050] Receive a shared memory protocol instruction; the instruction information contained in the shared memory protocol instruction includes: address information, operation code, source data;

[0051] According to the address information contained in the shared memory protocol instruction, the valid channel is judged to have the same protocol, and the valid channel is split according to the result of the same protocol judgment; the address information is used to indicate the address of the valid channel;

[0052] Reading data from the shared memory according to the splitting result of the valid channel; wherein the splitting result is associated with the number of times the data needs to be read from the shared memory;

[0053] A reduction operation is performed according to the data read from the shared memory, the source data and the operation code contained in the shared memory reduction instruction, and the result of the reduction operation is written back to the shared memory.

[0054] The above-mentioned shared memory access method, device, computer equipment, computer-readable storage medium and computer program product receive a shared memory protocol instruction; the instruction information contained in the shared memory protocol instruction includes: address information, operation code, source data; so that the shared memory can be accessed according to the shared memory protocol instruction to implement the protocol operation. According to the address information contained in the shared memory protocol instruction, the valid channel is judged to have the same protocol, and the valid channel is split according to the result of the same protocol judgment; the address information is used to indicate the address of the valid channel; so that the valid channels with the same address can be merged, and the valid channels with the same address are not read repeatedly. According to the split result of the valid channel, data is read from the shared memory; wherein the split result is associated with the number of times data needs to be read from the shared memory; thereby reducing the number of times the shared memory is read. The protocol operation is performed according to the data read from the shared memory, the source data and the operation code contained in the shared memory protocol instruction, and the result of the protocol operation is written back to the shared memory. This can automatically merge valid channels with the same address, reducing the number of read and write times of shared memory without adding hardware logic, significantly improving the efficiency of shared memory data protocol, and can be applied to shared memory with different configurations, different protocol opcodes, and support different data formats. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the drawings required for use in the embodiments of the present application or related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0056] Figure 1 An application environment diagram of a shared memory access method in one embodiment;

[0057] Figure 2 is a schematic diagram of a shared memory in an embodiment of the present application;

[0058] Figure 3 is a schematic diagram of a working group in an embodiment of the present application;

[0059] Figure 4 A schematic diagram of a flow chart of a method for accessing a shared memory in one embodiment;

[0060] Figure 5 A schematic diagram of a framework of a shared memory protocol operation in an embodiment of the present application;

[0061] FIG6 (a) is a schematic diagram of the same protocol determination process in an embodiment of the present application Figure 1 ;

[0062] FIG6( b ) is a schematic diagram of the same protocol determination process in an embodiment of the present application. Figure 2 ;

[0063] Figure 7 is a flow chart of a method for accessing a shared memory in another embodiment;

[0064] FIG8 (a) is a schematic diagram of information related to protocol operation in an embodiment of the present application Figure 1 ;

[0065] FIG8( b ) is a schematic diagram of information related to protocol operation in an embodiment of the present application Figure 2 ;

[0066] Fig. 9 This is a schematic diagram of the principle of effective channel splitting in one embodiment of the present application;

[0067] Fig.10 is a flow chart of a method for accessing a shared memory in yet another embodiment;

[0068] Fig.11 is a structural block diagram of a shared memory access device in one embodiment;

[0069] Fig.12 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0070] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0071] Before introducing in detail the shared memory access method provided by the present application, the following explanations are given for the computer terms used in the specification and claims of the present invention.

[0072] 1) Arithmetic logic unit: sends a read and / or write request and channel information of a shared memory to the shared memory control unit, for example, sends the address and / or data of a work item to the shared memory control unit.

[0073] 2) General register file: used to store source data required by the arithmetic logic unit, such as data returned by read operations received from the shared memory control unit for operation by the arithmetic logic unit.

[0074] 3) Shared memory control unit: used to process the read and / or write requests and channel information sent by the arithmetic logic operation unit, which includes a channel address continuity detection unit and a request splitting unit, controls the read operation and / or write operation of the shared memory, and is responsible for bypassing the data returned by the read request to the general register file; wherein, the channel address continuity detection unit also includes a processing mode selection unit.

[0075] 4) Shared memory: used to store the channel data of work items. The shared memory can be configured according to the number of memory banks m, the byte index of the memory bank width n, and the number of cache lines k.

[0076] The shared memory access method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or it can be placed on the cloud or other network servers. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptops, smart phones, tablets, Internet of Things devices and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart car devices, projection devices, etc. Portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The head-mounted device can be a virtual reality (VR) device, an augmented reality (AR) device, smart glasses, etc. The server 104 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services.

[0077] For example, Figure 2 Schematic diagram of a shared memory in an embodiment of the present application; Figure 2 As shown, the shared memory contains m banks, namely, banks Bank0, Bank1, Bank2, ..., Bankm-3, Bankm-2, Bankm-1, where m is an integer greater than 1. Each bank contains k cache lines, namely, cache lines Line0, Line1, Line2, Line3, ..., Linek-1, where k is also an integer greater than 1. Specifically, the storage width of Line0 in Bank0 is 2 n Bytes, where n is the byte index (such as Figure 2 Therefore, the entire shared memory (such as Figure 2 The size of 0, 1, 2...m-3, m-2, m-1...km-1) can be expressed as m*2 n*k Bytes.

[0078] It can be seen that the shared memory can be configured according to the number of memory banks m, the byte index of the memory bank width n, and the number of cache lines k. Therefore, the above parameters m, n, k are called configuration information of the shared memory.

[0079] For example, Figure 3 Schematic diagram of a working group in one embodiment of the present application; Figure 3 As shown in Figure 1, a work group can construct w threads (Wave1, Wave2, ..., Wave w ), each thread contains N lanes, and these N lanes usually execute the same instruction in parallel, forming a single instruction multiple data (SIMD) structure. The order of reduction is: starting from the first lane, the reduction operation is performed in sequence until the end of the Nth lane, and OP represents the operation. Exemplary atomic operations may include: atomic addition, atomic XOR, etc.

[0080] In the prior art, when threads of a work group perform reduction operations on data in a shared memory, the address of each channel in each thread must first be read from the shared memory, then reduced, and finally the reduction result is written to the shared memory through a write operation to the shared memory. In this way, multiple read and write operations will be performed on channels with the same address in the thread, resulting in a decrease in the performance of the shared memory reduction operation.

[0081] In response to the above problem, an embodiment of the present application detects the channel address of the workgroup thread and determines whether the addresses of each valid channel are exactly the same; when the addresses of the valid channels are the same, only one read operation and write operation is performed on the corresponding address in the shared memory, thereby reducing the number of read and write operations of the shared memory under the protocol operation and improving the protocol operation performance of the shared memory.

[0082] In an exemplary embodiment, Figure 4 As shown, a method for accessing a shared memory is provided. The method in this embodiment may include the following steps 401 to 402. Among them:

[0083] Step 401, receiving a shared memory protocol instruction.

[0084] The instruction information contained in the shared memory protocol instruction includes: address information, operation code, and source data.

[0085] For example, Figure 5 FIG. 1 is a schematic diagram of a framework of a shared memory protocol operation in an embodiment of the present application. Figure 5 As shown, the arithmetic logic unit sends a shared memory protocol instruction to the shared memory control unit. The instruction information contained in the shared memory protocol instruction includes: operation code, address information, source data, channel validity mask, etc. Figure 5 It can be seen that the instruction information contained in the shared memory protocol instruction sent by the arithmetic logic unit will be stored in the shared memory buffer. Then, the shared memory protocol instruction information control unit in the shared memory control unit reads the instruction information (such as address information, operation code, source data) from the shared memory buffer. The shared memory control unit is also used to execute instructions for the shared memory, such as executing shared memory read instructions, shared memory write instructions, and shared memory protocol instructions. Shared memory (SM) can be used to store workgroup data, where the shared memory parameters m, n, and k can be configured as needed.

[0086] Step 402, performing a same-protocol determination on valid channels according to the address information included in the shared memory protocol instruction, and splitting the valid channels according to the result of the same-protocol determination.

[0087] The address information is used to indicate the address of a valid channel.

[0088] Exemplary, combined Figure 5 It can be seen that the protocol instruction information control unit can perform the same protocol judgment on the valid channels according to the methods shown in Figure 6 (a) and Figure 6 (b) based on the address information contained in the shared memory protocol instruction. Figure 6 (a) shows the situation when the addresses of all valid channels are the same, and Figure 6 (b) shows a method for judging the same protocol of valid channels.

[0089] Exemplarily, the same protocol judgment refers to: obtaining the instruction information from the shared memory buffer through the protocol instruction information control unit; traversing to obtain the addresses of the valid channels, and performing pairwise XOR operations on the addresses of the valid channels. If the result of the pairwise XOR operation is not 0, it is determined that the addresses of the two valid channels are different; if the result of the pairwise XOR operation is 0, it is determined that the addresses of the two valid channels are the same.

[0090] For example, in conjunction with FIG6 (b), the address of the i-th valid channel (Addr i ) and the address of the i+1th valid channel (Addr i+1) performs an exclusive OR operation (XOR), where i is a natural number in [0, n-1] and n is the total number of valid channels. When the result of the exclusive OR operation between the address of the ith valid channel and the address of the i+1th valid channel is 0, the output result is "False". When the result of the exclusive OR operation between the address of the ith valid channel and the address of the i+1th valid channel is not 0, determine whether i+1 is equal to n. If so, the output result is "True". If not, let i increment by 1 and re-determine whether the result of the exclusive OR operation between the address of the ith valid channel and the address of the i+1th valid channel is 0, until the same protocol judgment of all channels is completed.

[0091] Exemplarily, valid channels with the same address are merged, and only valid channels with different addresses are retained; and the valid channels with different addresses are split according to the channel addresses corresponding to the valid channels.

[0092] In one possible case, combined with FIG6 (a), if the addresses of all valid channels are judged as 0 after the same protocol, that is, the addresses of all valid channels are the same, then the number of splits is determined to be 1, and the addresses of all valid channels are stored in a storage library.

[0093] In another possible case, in combination with FIG. 6 ( b ), if the addresses of valid channels are different, it is determined whether the addresses of the valid channels span different storage bins. If the addresses of the valid channels are all in one storage bin, the addresses of all valid channels are stored in one storage bin.

[0094] In another possible case, as shown in FIG. 6( b ), if the addresses of the valid channels are not all within one storage bin, the number of splits is determined according to the number of storage bins spanned by the addresses of the valid channels.

[0095] It should be noted that the configuration of the shared memory determines the number of storage banks and the width of each storage bank. Each storage bank corresponds to a split unit, and the number of split units corresponds to the number of times data needs to be read from the shared memory.

[0096] Step 403: Read data from the shared memory according to the splitting result of the valid channel.

[0097] The split result is associated with the number of times data needs to be read from the shared memory.

[0098] In this embodiment, combined with Figure 5 As shown, the protocol instruction information control unit sends a shared memory read request, and the shared memory read control module reads data from the shared memory according to the shared memory read request address.

[0099] Exemplarily, the number of times data needs to be read from the shared memory can be determined based on the splitting result of the valid channel, and the shared memory read request address corresponding to each read operation can be determined; and the corresponding data can be read from the shared memory based on the shared memory read request address.

[0100] Step 404, performing a reduction operation according to the data read from the shared memory, the source data and the operation code contained in the shared memory reduction instruction, and writing the result of the reduction operation back to the shared memory.

[0101] In this embodiment, combined with Figure 5 It can be seen that the reduction operation is performed based on the data read from the shared memory and the source data contained in the shared memory reduction instruction. Then, the result of the reduction operation is fed back to the shared memory write control module, and the shared memory write control module writes the result of the reduction operation back to the shared memory.

[0102] Exemplarily, the data read from the shared memory are matched with the valid channels indicated by each channel index to obtain data corresponding to each valid channel; according to the operation code contained in the shared memory protocol instruction, the data corresponding to each valid channel and the source data corresponding to each valid channel in the shared memory protocol instruction are reduced.

[0103] In the above-mentioned shared memory access method, by receiving a shared memory protocol instruction; the instruction information contained in the shared memory protocol instruction includes: address information, operation code, source data; thus, the shared memory can be accessed according to the shared memory protocol instruction to implement the protocol operation. According to the address information contained in the shared memory protocol instruction, the same protocol judgment is performed on the valid channel, and the valid channel is split according to the result of the same protocol judgment; the address information is used to indicate the address of the valid channel; thus, the valid channels with the same address can be merged, and the valid channels with the same address are not read repeatedly. According to the split result of the valid channel, data is read from the shared memory; wherein the split result is associated with the number of times data needs to be read from the shared memory; thereby reducing the number of times the shared memory is read. The protocol operation is performed according to the data read from the shared memory, the source data and the operation code contained in the shared memory protocol instruction, and the result of the protocol operation is written back to the shared memory. This can automatically merge valid channels with the same address, reducing the number of read and write times of shared memory without adding hardware logic, significantly improving the efficiency of shared memory data protocol, and can be applied to shared memory with different configurations, different protocol opcodes, and support different data formats.

[0104] In another exemplary embodiment, Figure 7As shown, a method for accessing a shared memory is provided. The method in this embodiment may include steps 701 to 705. Among them:

[0105] Step 701, receiving a shared memory protocol instruction.

[0106] Step 702: Perform same-protocol determination on valid channels according to the address information included in the shared memory protocol instruction, and split the valid channels according to the result of the same-protocol determination.

[0107] Step 703: Read data from the shared memory according to the splitting result of the valid channel.

[0108] Step 704, performing a reduction operation according to the data read from the shared memory, the source data and the operation code contained in the shared memory reduction instruction, and writing the result of the reduction operation back to the shared memory.

[0109] In this embodiment, the specific implementation process and technical effects of steps 701 to 704 are shown in Figure 4 The descriptions of steps 401 to 404 in the illustrated method embodiment are not repeated here.

[0110] Step 705: When the identification information of whether to write back the reduction result is not empty, the result of the reduction operation is written back to the general file register.

[0111] In this embodiment, the general file register is used to store the result returned by the shared memory protocol instruction operation.

[0112] Exemplarily, the instruction information included in the shared memory protocol instruction also includes: the number of valid channels, identification information of whether to write back the protocol result, and the data format supported by the protocol. When the identification information of whether to write back the protocol result is NULL, the result of the protocol operation may not be written into the general file register.

[0113] For example, assume that the shared memory configuration (m, n, k) = (16, 2, 1024), that is, the number of banks is 16 and the width of the bank is 2. 2 bits, and the number of storage lines is 1024. Based on this, the following shared memory protocol instructions are generated:

[0114] (Pn)SM_REDU Dest, Addr, Src_Data2, Src_Data1, OpCode, data_fmt

[0115] Where: Pn represents the prediction file register, which is used to store the valid channel of the shared memory protocol instruction; Dest represents whether the result of the protocol operation needs to be written back to the general file register, and the value of Dest can be empty (that is, the result of the protocol operation is not written back); Addr represents the address of the valid channel of the shared memory protocol instruction; Src_Data1 and Src_Data2 represent the source data of the shared memory protocol instruction; OpCode represents the operation code of the shared memory protocol instruction, such as: addition operation, minimum operation, maximum operation, etc.; data_fmt represents the data format supported by the protocol operation, such as 32 bits, 64 bits, etc.

[0116] For example, assume that there are two shared memory protocol instructions as follows:

[0117] (P1)SM_REDU NULL,R10, VOID, R16, ADD, 32bits

[0118] (P1)SM_REDU R26, R20, VOID, R18, ADD, 32bits

[0119] Among them: P1 = 0x91010505 (there are 32 valid channels), the addresses of the valid channels of the shared memory protocol instructions are R10 and R20, and the source data are R16 and R18; and both perform 32-bit addition operations.

[0120] For example, as shown in Figure 8 (a), the channel indexes are: 0, 2, 8, 10, 16, 24, 28, 31, and the channel addresses are all 4, that is, the addresses of the valid channels are all the same, the source data of each channel are: 1, 2, 3, 4, 5, 6, 7, 8, and the reduced operation code is: add operation (ADD). Since the addresses of all valid channels are the same, the valid channels can be merged, the number of splits is 1, and only one read operation and one write operation are performed on the shared memory.

[0121] For example, as shown in FIG8( b ), the channel indexes are: 0, 2, 8, 10, 16, 24, 28, 31, and the channel addresses are 0, 8, 12, 16, 12, 16, 20, 24, that is, the addresses of the valid channels are all the same, and the source data of the channels are: 1, 2, 3, 4, 5, 6, 7, 8, and the protocol operation code is: add operation. Since the addresses of the channels are not all the same, the splitting mode of the valid channels can be determined according to whether the addresses of the valid channels span multiple storage banks. For example, Fig. 9 As shown, the processing request mode 16P1C can be used to split the valid channel, thereby obtaining the following two requests:

[0122] Request 1: includes valid channels 0, 2, 8, and 10;

[0123] Request 2: includes valid channels 16, 24, 28, 31.

[0124] The 16 in the processing request mode 16P1C indicates that the data format is 16 bits, and P1C indicates a non-integrated mode across multiple storage banks.

[0125] In this embodiment, by splitting the request into two times, only two read operations and two write operations are performed when performing a shared memory protocol operation.

[0126] In yet another exemplary embodiment, Fig.10 As shown, a method for accessing a shared memory is provided, which may include the following steps:

[0127] Step S1: Let i=0, and obtain the first valid channel Lanei;

[0128] Step S2: Point to channel Lanei;

[0129] Step S3: Obtain channel data of channel Lanei (denoted as Lane[i]) and source data of channel Lanei;

[0130] Step S4: performing a reduction operation according to the channel data and source data of the channel Lanei to obtain a result of the reduction operation;

[0131] Step S5: updating the channel data Lane[i] of the channel Lanei according to the result of the reduction operation;

[0132] Step S6: Determine whether the value of i+1 is equal to n; if so, write the result of the reduction operation back to the shared memory; if not, set the channel data of channel Lanei+1 to Lane[i], and increment i by 1, then return to execute step S2.

[0133] Exemplarily, if the identification information of whether to write back the reduction result in the shared memory reduction instruction is not empty, the reduction operation result is returned to the general register file.

[0134] In this embodiment, by detecting the address of the channel in the work group thread, the channels with the same address are merged, so that only one shared memory read operation and one shared memory write operation are required for the same channel address, thereby greatly improving the efficiency of the shared memory data protocol. In addition, the entire protocol operation does not need to add additional hardware logic, and only needs to perform the same protocol judgment on the effective channel address of the shared memory protocol instruction (two-by-two address XOR operation), and the protocol efficiency can be improved without complex hardware logic, and the cost is low. The method in this embodiment can be applied to different shared memory configurations, and can adapt to different protocol opcodes, and support different data formats.

[0135] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0136] Based on the same inventive concept, the embodiment of the present application also provides a shared memory access device for implementing the shared memory access method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in the one or more shared memory access device embodiments provided below can refer to the limitations on the shared memory access method above, and will not be repeated here.

[0137] In an exemplary embodiment, Fig.11 As shown, a shared memory access device is provided, including: a receiving module 1101, a judging module 1102, a reading module 1103 and a protocol module 1104, wherein:

[0138] The receiving module 1101 is used to receive a shared memory protocol instruction; the instruction information contained in the shared memory protocol instruction includes: address information, operation code, source data;

[0139] The judgment module 1102 is used to perform the same protocol judgment on the valid channels according to the address information contained in the shared memory protocol instruction, and split the valid channels according to the result of the same protocol judgment; the address information is used to indicate the address of the valid channel;

[0140] A reading module 1103 is used to read data from the shared memory according to the splitting result of the valid channel; wherein the splitting result is associated with the number of times the data needs to be read from the shared memory;

[0141] The reduction module 1104 is used to perform a reduction operation according to the data read from the shared memory, the source data and the operation code contained in the shared memory reduction instruction, and write the result of the reduction operation back to the shared memory.

[0142] In one embodiment, the receiving module 1101 is specifically used for:

[0143] receiving a shared memory protocol instruction sent by an arithmetic logic operation unit;

[0144] The instruction information contained in the shared memory protocol instruction is stored in the shared memory buffer.

[0145] In one embodiment, the determination module 1102 is specifically configured to

[0146] Obtaining the instruction information from the shared memory buffer through a protocol instruction information control unit;

[0147] Traverse to obtain the addresses of valid channels, and perform pairwise XOR operations on the addresses of valid channels. If the result of the pairwise XOR operation is not 0, it is determined that the addresses of the two valid channels are different; if the result of the pairwise XOR operation is 0, it is determined that the addresses of the two valid channels are the same;

[0148] Merge valid channels with the same address and only keep valid channels with different addresses;

[0149] The valid channels with different addresses are split according to the channel addresses corresponding to the valid channels.

[0150] In one embodiment, the traversing to obtain addresses of valid channels and performing pairwise XOR operations on the addresses of valid channels include:

[0151] Perform an XOR operation on the address of the i-th valid channel and the address of the i+1-th valid channel, where i is a natural number in [0, n-1] and n is the total number of valid channels.

[0152] In one embodiment, splitting the valid channels according to the channel addresses corresponding to the valid channels with different addresses includes:

[0153] If the addresses of all valid channels are the same, the number of splits is determined to be 1, and the addresses of all valid channels are stored in a storage library;

[0154] If the addresses of the valid channels are different, it is determined whether the addresses of the valid channels span different storage bins. If the addresses of the valid channels are all in one storage bin, the addresses of all the valid channels are stored in one storage bin. If the addresses of the valid channels are not all in one storage bin, the number of splits is determined according to the number of storage bins spanned by the addresses of the valid channels.

[0155] The configuration of the shared memory determines the number of storage banks and the width of each storage bank. Each storage bank corresponds to a split unit, and the number of split units corresponds to the number of times data needs to be read from the shared memory.

[0156] In one embodiment, the reading module 1103 is specifically used for:

[0157] Determine the number of times data needs to be read from the shared memory according to the split result of the valid channel, and determine the shared memory read request address corresponding to each read operation;

[0158] According to the shared memory read request address, corresponding data is read from the shared memory.

[0159] In one embodiment, the protocol module 1104 is specifically used to:

[0160] Match the data read from the shared memory with the valid channels indicated by each channel index to obtain the data corresponding to each valid channel;

[0161] According to the operation code included in the shared memory protocol instruction, a protocol operation is performed on the data corresponding to each valid channel and the source data corresponding to each valid channel in the shared memory protocol instruction.

[0162] In one embodiment, the instruction information included in the shared memory protocol instruction also includes: the number of valid channels, identification information of whether to write back the protocol result, and the data format supported by the protocol.

[0163] In one of the embodiments, the reduction module 1104 in the above device is further configured to write the result of the reduction operation back to the general file register when the identification information of whether to write back the reduction result is not empty.

[0164] Each module in the above-mentioned shared memory access device can be implemented in whole or in part by software, hardware or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute operations corresponding to each module.

[0165] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Fig.12As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a shared memory access method is implemented.

[0166] Those skilled in the art will understand that Fig.12 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0167] In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.

[0168] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0169] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0170] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0171] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited to this.

[0172] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0173] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A method for accessing a shared memory, characterized in that: The method comprises: Receive a shared memory protocol instruction; the instruction information contained in the shared memory protocol instruction includes: address information, operation code, source data; According to the address information contained in the shared memory protocol instruction, the valid channel is judged to have the same protocol, and the valid channel is split according to the result of the same protocol judgment; the address information is used to indicate the address of the valid channel; Reading data from the shared memory according to the splitting result of the valid channel; wherein the splitting result is associated with the number of times the data needs to be read from the shared memory; A reduction operation is performed according to the data read from the shared memory, the source data and the operation code contained in the shared memory reduction instruction, and the result of the reduction operation is written back to the shared memory.

2. The method according to claim 1, characterized in that The receiving of the shared memory protocol instruction comprises: receiving a shared memory protocol instruction sent by an arithmetic logic operation unit; The instruction information contained in the shared memory protocol instruction is stored in the shared memory buffer.

3. The method according to claim 2, characterized in that The step of performing same-rule judgment on valid channels according to the address information included in the shared memory protocol instruction, and splitting the valid channels according to the result of the same-rule judgment, includes: Obtaining the instruction information from the shared memory buffer through a protocol instruction information control unit; Traverse to obtain the addresses of valid channels, and perform pairwise XOR operations on the addresses of valid channels. If the result of the pairwise XOR operation is not 0, it is determined that the addresses of the two valid channels are different; if the result of the pairwise XOR operation is 0, it is determined that the addresses of the two valid channels are the same; Merge valid channels with the same address and only keep valid channels with different addresses; The valid channels with different addresses are split according to the channel addresses corresponding to the valid channels.

4. The method according to claim 3, characterized in that The traversal obtains the addresses of valid channels and performs pairwise XOR operations on the addresses of valid channels, including: Perform an XOR operation on the address of the i-th valid channel and the address of the i+1-th valid channel, where i is a natural number in [0, n-1] and n is the total number of valid channels.

5. The method according to claim 3, characterized in that: The splitting of the valid channels according to the channel addresses corresponding to the valid channels with different addresses includes: If the addresses of all valid channels are the same, the number of splits is determined to be 1, and the addresses of all valid channels are stored in a storage library; If the addresses of the valid channels are different, it is determined whether the addresses of the valid channels span different storage bins. If the addresses of the valid channels are all in one storage bin, the addresses of all the valid channels are stored in one storage bin. If the addresses of the valid channels are not all in one storage bin, the number of splits is determined according to the number of storage bins spanned by the addresses of the valid channels. The configuration of the shared memory determines the number of storage banks and the width of each storage bank. Each storage bank corresponds to a split unit, and the number of split units corresponds to the number of times data needs to be read from the shared memory.

6. The method according to any one of claims 1 to 5, characterized in that: The step of reading data from a shared memory according to the splitting result of the valid channel comprises: Determine the number of times data needs to be read from the shared memory according to the split result of the valid channel, and determine the shared memory read request address corresponding to each read operation; According to the shared memory read request address, corresponding data is read from the shared memory.

7. The method according to any one of claims 1 to 5, characterized in that: The performing of the reduction operation according to the data read from the shared memory, the source data and the operation code contained in the shared memory reduction instruction includes: Match the data read from the shared memory with the valid channels indicated by each channel index to obtain the data corresponding to each valid channel; According to the operation code included in the shared memory protocol instruction, a protocol operation is performed on the data corresponding to each valid channel and the source data corresponding to each valid channel in the shared memory protocol instruction.

8. The method according to any one of claims 1 to 5, characterized in that: The instruction information contained in the shared memory protocol instruction also includes: the number of valid channels, identification information of whether to write back the protocol result, and the data format supported by the protocol.

9. The method according to claim 8, characterized in that After writing the result of the reduction operation back to the shared memory, the method further comprises: When the identification information of whether to write back the reduction result is not empty, the result of the reduction operation is written back to the general file register.

10. A shared memory access device, characterized in that: The device comprises: A receiving module, used for receiving a shared memory protocol instruction; the instruction information contained in the shared memory protocol instruction includes: address information, operation code, source data; A judgment module, used for performing same-rule judgment on valid channels according to the address information contained in the shared memory protocol instruction, and splitting the valid channels according to the result of the same-rule judgment; the address information is used to indicate the address of the valid channel; A reading module, used for reading data from the shared memory according to the splitting result of the valid channel; wherein the splitting result is associated with the number of times the data needs to be read from the shared memory; The protocol module is used to perform a protocol operation according to the data read from the shared memory, the source data and the operation code contained in the shared memory protocol instruction, and write the result of the protocol operation back to the shared memory.

11. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 9 are implemented.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.

13. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.

Citation Information

Cited By

  • Method and device for realizing reduction algorithm

    CN115345290A

  • A method and apparatus for implementing a reduction algorithm

    CN115345290B