Data Reading Method, Device, Electronic Device, and Computer-Readable Storage Medium
By setting at least two address memory to write multiple read addresses in parallel for each storage unit, the read conflict problem is solved, and the parallel processing efficiency of the computing device and the utilization rate of the transmission device are improved.
Patent Information
- Application Number
- CN202510180229.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-19
AI Technical Summary
The prior art is prone to read conflicts when processing multiple read addresses, resulting in backpressure of the data input terminal, reducing the parallel processing efficiency of the computing device and the bandwidth of the storage device.
At least two address memories are provided for each storage unit, and multiple read addresses are written in parallel to the corresponding address memories, thereby reducing the possibility of read conflicts and reducing the number of transmissions through data merging.
It effectively solves the problem of read conflict, improves the parallel processing efficiency of computing devices and the utilization rate of transmission devices, and improves the overall performance of data reading.
Smart Images

Figure CN119645321B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to a data reading method, apparatus, electronic device, and computer-readable storage medium. Background Art
[0002] As the computing power of computing devices becomes stronger and the amount of data to be processed becomes larger, the central processing unit (CPU) and coprocessors have been developing rapidly. The most common coprocessor is the graphics processing unit (GPU), which mainly processes computations related to graphics display. Later, it developed to using the GPU to process general computing tasks, thus emerging the general-purpose computing on graphics processing units (GPGPU). The GPGPU chip can not only provide powerful graphics processing capabilities, but also be used to execute highly parallel scientific computing, data analysis, machine learning, and other computationally intensive tasks. The GPGPU chip is usually equipped with high-speed video memory and has a higher memory bandwidth than the CPU, which helps to accelerate data-intensive applications. Summary of the Invention
[0003] At least one embodiment of the present disclosure provides a data reading method for reading data from M storage units, where each of the storage units corresponds to a group of address memories, and each group of address memories includes at least two address memories. The method includes: receiving a read request, where the read request includes N groups of read addresses respectively corresponding to N of the M storage units, and each of the N groups of read addresses includes one or more read addresses; writing the N groups of read addresses into the N groups of address memories corresponding to the N storage units respectively. Among the N groups of read addresses, the first group of read addresses corresponds to the first storage unit among the N storage units. When the first group of read addresses includes at least two read addresses and the number of addresses in the first group of read addresses is not greater than the number of memories in the first group of address memories corresponding to the first storage unit, the first group of read addresses is written into the first group of address memories in parallel; reading sub-data corresponding to the read addresses in the N groups of address memories from the N storage units to obtain N groups of sub-data respectively corresponding to the N groups of read addresses; merging the N groups of sub-data to obtain target data corresponding to the read request, and transmitting the target data to a target device through a data transmission device; where M is an integer greater than 1, and N is an integer greater than or equal to 1 and less than M.
[0004] For example, in the data reading method provided by at least one example of the foregoing embodiments of the present disclosure, the first set of read addresses includes k read addresses, where k is an integer greater than or equal to 2; writing the N sets of read addresses into the N sets of address memories corresponding to the N storage units respectively includes: when the number of memories in the first set of address memories is greater than or equal to k, writing the k read addresses into k address memories in the first set of address memories in parallel.
[0005] For example, in the data reading method provided by at least one example of the foregoing embodiments of the present disclosure, writing the N sets of read addresses into the N sets of address memories corresponding to the N storage units respectively includes: when the number of addresses included in each set of the N sets of read addresses is not greater than the number of memories included in a corresponding set of address memories, writing the N sets of read addresses into the N sets of address memories in parallel.
[0006] For example, in the data reading method provided by at least one example of the foregoing embodiments of the present disclosure, reading sub-data corresponding to the read addresses in the N sets of address memories from the N storage units to obtain N sets of sub-data corresponding to the N sets of read addresses respectively includes: in each clock cycle, based on an arbitration rule, selecting one address memory from the first set of address memories; reading the sub-data corresponding to the read address in the selected address memory from the first storage unit.
[0007] For example, in the data reading method provided by at least one example of the foregoing embodiments of the present disclosure, each of the address memories includes a plurality of storage locations, and the width of each storage location in the plurality of storage locations corresponds to the data bit width of one of the read addresses; wherein, in each clock cycle, based on an arbitration rule, selecting one address memory from the first set of address memories includes: selecting one address memory from at least two address memories in the first set of address memories based on the polling order of at least two address memories in the first set of address memories; wherein, reading the sub-data corresponding to the read address in the selected address memory from the first storage unit includes: reading the sub-data corresponding to the first read address in the selected address memory from the first storage unit.
[0008] For example, in the data reading method provided by at least one example of the foregoing embodiments of the present disclosure, it further includes: recording the correspondence between each read address in the read request and the N sets of address memories to form record information.
[0009] For example, in the data reading method provided by at least one example of the above embodiments of the present disclosure, each of the storage units corresponds to a group of data memories, each group of data memories includes at least two data memories, and a group of address memories and a group of data memories corresponding to each of the storage units correspond to each other; wherein, reading sub-data corresponding to the read addresses in the N groups of address memories from the N storage units to obtain N groups of sub-data corresponding to the N groups of read addresses respectively, includes: for each of the storage units, after reading the sub-data corresponding to one read address each time, selecting one data memory from the corresponding group of data memories, and writing the read sub-data into the selected data memory; based on the record information, reading the sub-data corresponding to the read request from the N groups of data memories to obtain the N groups of sub-data.
[0010] For example, in the data reading method provided by at least one example of the above embodiments of the present disclosure, based on the record information, reading the sub-data corresponding to the read request from the N groups of data memories to obtain the N groups of sub-data, includes: when the number of addresses included in each group of the N groups of read addresses is not greater than the number of memories included in the corresponding group of address memories, based on the record information, simultaneously reading the sub-data corresponding to the read request from the N groups of data memories to obtain the N groups of sub-data.
[0011] For example, in the data reading method provided by at least one example of the above embodiments of the present disclosure, writing the N groups of read addresses into the N groups of address memories corresponding to the N storage units respectively, includes: when the number of addresses in the first group of read addresses is greater than the number of memories in the first group of address memories corresponding to the first storage unit, writing the first group of read addresses in multiple times. Among them, based on the record information, reading the sub-data corresponding to the read request from the N groups of data memories, includes: when there is at least one group of read addresses among the N groups of read addresses whose number of included addresses is greater than the number of memories included in the corresponding group of address memories, reading the sub-data corresponding to the read request from the N groups of data memories in multiple times. Among them, merging the N groups of sub-data to obtain the target data corresponding to the read request, includes: storing the sub-data read from the N groups of data memories each time into the merge memory until the sub-data corresponding to the read request are all stored in the merge memory, and merging the sub-data in the merge memory to obtain the target data corresponding to the read request.
[0012] At least one embodiment of the present disclosure provides a data reading device, including M groups of address memories, an address processing unit, M reading units, and a data processing unit. The M groups of address memories correspond to M storage units, and each group of address memories includes at least two address memories. The address processing unit is connected to the M groups of address memories, and the M reading units correspond to the M storage units. Among them, the address processing unit is configured to: receive a reading request, where the reading request includes N groups of read addresses respectively corresponding to N storage units among the M storage units, and each group of the N groups of read addresses includes one or more read addresses; write the N groups of read addresses into the N groups of address memories corresponding to the N storage units respectively, where the N groups of read addresses include a first group of read addresses corresponding to a first storage unit among the N storage units, and when the first group of read addresses includes at least two read addresses and the number of addresses in the first group of read addresses is not greater than the number of memories in the first group of address memories corresponding to the first storage unit, the first group of read addresses is written into the first group of address memories in parallel. Among them, the M reading units are configured to read sub-data corresponding to the read addresses in the N groups of address memories from the N storage units to obtain N groups of sub-data respectively corresponding to the N groups of read addresses. Among them, the data processing unit is configured to merge the N groups of sub-data to obtain target data corresponding to the reading request, and transmit the target data to a target device through a data transmission unit.
[0013] For example, in the data reading device provided in at least one example of the above embodiment of the present disclosure, the M reading units include a first reading unit corresponding to the first storage unit, and the first reading unit is further configured to: in each clock cycle, based on an arbitration rule, select one address memory from the first group of address memories; read sub-data corresponding to the read address in the selected address memory from the first storage unit.
[0014] For example, in the data reading device provided in at least one example of the above embodiment of the present disclosure, it further includes: a recording unit, configured to record the correspondence between each read address in the reading request and the N groups of address memories to form recording information.
[0015] For example, in the data reading device provided in at least one example of the foregoing embodiments of the present disclosure, it further includes M groups of data memories and a transmission unit. The M groups of data memories correspond to M storage units, and each group of data memories includes at least two data memories; the transmission unit is connected between the M groups of data memories and the data processing unit; wherein, each of the reading units is further configured to: each time after reading sub-data corresponding to a read address from the corresponding storage unit, select one data memory from the corresponding group of data memories, and write the read sub-data into the selected data memory; wherein, the transmission unit is configured to read sub-data corresponding to the reading request from the N groups of data memories based on the record information, and send the sub-data to the data processing unit.
[0016] At least one embodiment of the present disclosure provides an electronic device, including a processor; a memory storing one or more computer program modules; wherein, the one or more computer program modules are configured to be executed by the processor to implement the data reading method provided in any embodiment of the present disclosure.
[0017] At least one embodiment of the present disclosure provides a computer-readable storage medium storing non-transitory computer-readable instructions, which can implement the data reading method provided in any embodiment of the present disclosure when executed by a computer. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly introduced below. Obviously, the drawings in the following description only relate to some embodiments of the present disclosure and do not limit the present disclosure.
[0019] Figure 1A It is a schematic diagram of a storage unit organization method;
[0020] Figure 1B It is a schematic diagram of another storage unit organization method;
[0021] Figure 2 It shows a flowchart of a data reading method provided in at least one embodiment of the present disclosure;
[0022] Figure 3 It shows a schematic diagram of a data reading process provided in at least one embodiment of the present disclosure;
[0023] Figure 4 It shows a flowchart of reading sub-data corresponding to N groups of read addresses provided in at least one embodiment of the present disclosure;
[0024] Figure 5A and 5BAnother schematic diagram showing the data reading process provided by at least one embodiment of the present disclosure;
[0025] Figure 6 A schematic block diagram showing a data reading device provided by at least one embodiment of the present disclosure;
[0026] Figure 7 A schematic block diagram showing an electronic device provided by at least one embodiment of the present disclosure;
[0027] Figure 8 A schematic block diagram showing another electronic device provided by at least one embodiment of the present disclosure; and
[0028] Figure 9 A schematic diagram showing a computer-readable storage medium provided by at least one embodiment of the present disclosure. Detailed implementation manners
[0029] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present disclosure. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.
[0030] Unless otherwise defined, the technical terms or scientific terms used in the present disclosure shall have the ordinary meanings understood by those of ordinary skill in the art to which the present disclosure pertains. The "first", "second", and similar terms used in the present disclosure do not denote any order, quantity, or importance, but are only used to distinguish different components. Similarly, the terms such as "a", "an", or "the" do not denote a quantity limitation, but mean that there is at least one. The terms such as "include" or "comprise" mean that the elements or items appearing before the term cover the elements or items listed after the term and their equivalents, without excluding other elements or items. The terms such as "connect" or "couple" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left", "right", etc. are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0031] Inside the processor, there is shared memory which allows multiple processes to access the same memory area, enabling fast data exchange between these processes. Shared memory can efficiently transfer large amounts of data. Since data does not need to be copied between different processes, as a very important storage unit within a processor (such as a GPU), the efficiency of reading memory directly affects the data processing efficiency of the entire processor chip. Usually, in order to achieve high-bandwidth simultaneous access to memory, the shared memory is divided into equally sized storage units that can be accessed in parallel. For example, multiple independent storage units can be organized in a multi-banking manner, where each storage unit can be a bank, or it can be a bank group containing multiple banks. For shared memory, the divided banks can also be referred to as memory blocks. When accessing multiple addresses, access operations can be performed simultaneously on multiple independent storage units. However, once there is a simultaneous read of the same storage unit, an access conflict will occur.
[0032] Figure 1A It is a schematic diagram of an organization method of a storage unit. Figure 1B It is a schematic diagram of another organization method of a storage unit.
[0033] For example, as Figure 1A and Figure 1B shown, the internal storage device 110 can be the internal storage device of a computing device (such as a CPU, GPU, GPGPU, or NPU, etc.). The internal storage device 110 communicates with the request device 130 through the bus 120. For example, the internal storage device 110 can be various types of memory, cache, etc., and the request device 130 can be devices such as a processor, a memory, etc. The embodiments of the present disclosure do not limit this.
[0034] For example, as Figure 1A shown, taking a bank as a storage unit as an example, the internal storage device 110 includes N banks (bank 0, bank 1... bank N - 1), where N is an integer greater than 1. For example, when the request device 130 in the computing device wants to perform a data read access to multiple banks, it will send a corresponding read request to the bus 120 to send the read request to the internal storage device 110 through the bus 120. The read request contains multiple read addresses, and the read addresses correspond to the banks to be accessed in the internal storage device 110; after reading data from the bank, the read data is transmitted to the request device through the bus.
[0035] For example, as Figure 1BAs shown in the figure, taking the bank group as the storage unit as an example, the internal storage device 110 includes N bank groups (bank group 0, bank group 1... bank group N-1), and each bank group includes n banks (bank 0, bank 1... bank n-1), where both N and n are integers greater than 1. For example, when the computing device needs to access or store data in multiple bank groups, Figure 1B And Figure 1A The difference in the memory access process is that the read address corresponds to the bank group to be accessed in the internal storage device 110.
[0036] For example, in the Figure 1A And Figure 1B In the example, under normal circumstances, the read address corresponds one-to-one with the multiple banks / bank groups to be accessed; when two or more read addresses correspond to the same bank / bank group, a read conflict will occur, or it is also called a bank conflict, that is, it is impossible to simultaneously read the data corresponding to two or more read addresses from the same bank / bank group.
[0037] For example, when a read conflict occurs between read address 1 and read address 2 in one group of read addresses, in the current clock cycle, it is only possible to parallelly read the data of the corresponding bank / bank group based on read address 1 and other read addresses except read address 2, while for read address 2, it is only possible to separately read the data of the corresponding bank / bank group in the next clock cycle. In this way, on the one hand, it will generate backpressure on the bus 120, affecting the overall performance of the hardware device; on the other hand, since an additional clock cycle is added to separately read read address 2, the bandwidth of the storage device is reduced by half. For the case where read conflicts frequently occur, the bandwidth of the storage device will be greatly reduced.
[0038] On the transmission link of the computing device, write conflicts are often encountered. As the number of processing cores in the computing device increases, the memory access to the storage units in the storage device also increases accordingly, thereby increasing the possibility of write conflicts, reducing the bandwidth of the storage device, greatly increasing the data memory access time, and reducing the parallel processing efficiency of the computing device.
[0039] In some cases, a suitable swizzle algorithm can be used to handle read conflicts. For example, for different read addresses, different swizzle patterns are used to reassign the addresses corresponding to the storage units in the storage device through hashing and other methods, so that multiple read addresses with read conflicts are re-corresponded to multiple different storage units. This situation requires software participation, greatly increasing the software complexity.
[0040] In addition, for multiple read addresses with read conflicts, if they are processed over multiple clock cycles, the read data will also be transmitted back to the requesting device over multiple clock cycles. In this way, the bus needs to perform multiple data transmissions, resulting in low efficiency and increased bus load.
[0041] At least one embodiment of the present disclosure provides a data reading method, a data reading device, an electronic device, and a computer-readable storage medium. The data reading method includes: receiving a reading request, where the reading request includes N groups of read addresses respectively corresponding to N of the M storage units, and each group of the N groups of read addresses includes one or more read addresses; writing the N groups of read addresses into N groups of address memories corresponding to the N storage units respectively, where the N groups of read addresses include a first group of read addresses corresponding to a first storage unit among the N storage units, and when the first group of read addresses includes at least two read addresses and the number of addresses in the first group of read addresses is not greater than the number of memories in the first group of address memories corresponding to the first storage unit, the first group of read addresses is written into the first group of address memories in parallel; reading sub-data corresponding to the read addresses in the N groups of address memories from the N storage units to obtain N groups of sub-data respectively corresponding to the N groups of read addresses; merging the N groups of sub-data to obtain target data corresponding to the reading request, and transmitting the target data to a target device through a data transmission device; where M is an integer greater than 1, and N is an integer greater than or equal to 1 and less than M.
[0042] In this data reading method, by setting at least two address memories for each storage unit, multiple read addresses with read conflicts can be first written into the address memories corresponding to the storage units, reducing the backpressure on the data input end, solving the problem of read conflicts at the hardware level, and improving the parallel processing efficiency of the computing device. Moreover, by merging the return data corresponding to the conflicting read addresses, the number of transmissions of the transmission device is reduced, and the utilization rate of the transmission device is improved.
[0043] Figure 2 The flowchart of a data reading method provided by at least one embodiment of the present disclosure is shown.
[0044] As Figure 2 shown, this data reading method is used to read data from M storage units, each storage unit corresponds to a group of address memories, and each group of address memories includes at least two address memories. This method may include steps S210 to S240.
[0045] Step S210: Receive a read request, where the read request includes N groups of read addresses respectively corresponding to N of the M storage units, and each group of the N groups of read addresses includes one or more read addresses.
[0046] Step S220: Write the N groups of read addresses into N groups of address memories respectively corresponding to the N storage units. The N groups of read addresses include a first group of read addresses corresponding to a first storage unit among the N storage units. When the first group of read addresses includes at least two read addresses and the number of addresses in the first group of read addresses is not greater than the number of memories in the first group of address memories corresponding to the first storage unit, the first group of read addresses is written into the first group of address memories in parallel;
[0047] Step S230: Read sub-data corresponding to the read addresses in the N groups of address memories from the N storage units to obtain N groups of sub-data respectively corresponding to the N groups of read addresses;
[0048] Step S240: Combine the N groups of sub-data to obtain target data corresponding to the read request, and transmit the target data to the target device through a data transmission device.
[0049] For example, M is an integer greater than 1, and N is an integer greater than or equal to 1 and less than M.
[0050] For example, the M storage units may be multiple storage units obtained by partitioning the shared memory of a processor, and the storage units may be the above-mentioned banks or bank groups. The embodiments of the present disclosure do not limit the types and numbers of storage units.
[0051] Figure 3 The figure shows a schematic diagram of a data reading process provided by at least one embodiment of the present disclosure.
[0052] Such as Figure 3As shown in the figure, taking the storage unit as a bank group as an example, each storage unit includes multiple banks. For example, bank0 to bank3 form bank group 0, bank4 to bank7 form bank group 1, …, bank28 to bank31 form bank group 7. For each bank group, a set of address memories is provided. The number of address memories included in each set of address memories is not limited in the embodiments of the present disclosure. Each set of address memories may include two or more address memories. Hereinafter, an example in which each set of address memories includes two address memories will be described. For example, for bank group 0, address memories A0 to A1 are provided; for bank group 1, address memories A2 to A3 are provided; for bank group 2, address memories A4 to A5 are provided; for bank group 3, address memories A6 to A7 are provided, …, for bank group 7, address memories A14 to A15 are provided. For example, each address memory can be implemented as a device for storage such as a register.
[0053] For example, in step S210, the address processing unit can receive a read request and can parse to obtain multiple read addresses in the read request. After obtaining the multiple read addresses, the address processing unit can send each read address to the corresponding address memory through a first selector (MUX). Each read address can correspond to a bank group. In some cases, there are two or more read addresses corresponding to the same bank group. For example, the read request includes read addresses r0 to r4 (only r0 to r2 are shown in the figure), where read addresses r0 and r1 correspond to bank group 0, and read addresses r2 to r4 correspond to bank group 1 to bank group 3 respectively. In this case, one or more read addresses corresponding to each bank group can be used as a set of read addresses, and each set of read addresses can include one or more read addresses. For example, read addresses r0 and r1 can be used as a set of read addresses, read address r2 can be regarded as a set of read addresses, read address r3 can be regarded as a set of read addresses, and read address r4 can be regarded as a set of read addresses.
[0054] For example, the first storage unit can be any one of the storage units corresponding to multiple (two or more) read addresses in the read request. In the above example, bank group 0 corresponds to two read addresses r0 and r1, and the first storage unit can be bank group 0. For any storage unit corresponding to multiple read addresses, the operations for the first storage unit can be referred to.
[0055] For example, a first group of address memories is provided for a first storage unit. The first storage unit corresponds to a first group of read addresses in a read request. The first group of read addresses includes k read addresses, where k is an integer greater than or equal to 2. In step S220, when the number of memories in the first group of address memories is greater than or equal to k, the k read addresses are written into k address memories in the first group of address memories in parallel. Writing the k read addresses into k address memories in parallel can be understood as writing the k read addresses into k address memories respectively within the same clock cycle. For example, the first storage unit is bank group0, the first group of address memories is address memories A0 and A1, and the first group of read addresses is read addresses r0 and r1. In this example, k is 2, and the number of address memories included in the first group of address memories is equal to k. Therefore, the read addresses r0 and r1 can be written into address memories A0 and A1 in parallel. For example, within the same clock cycle, the read address r0 is written into address memory A0, and the read address r1 is written into address memory A1. In this way, when there is a read conflict, the conflicting read addresses can be written into the address memories simultaneously, ensuring that the read addresses in the read request do not backpressure the input end.
[0056] For example, in step S220, when the number of addresses included in each of the N groups of read addresses is not greater than the number of memories included in the corresponding group of address memories, the N groups of read addresses are written into the N groups of address memories in parallel. For example, bank group 0 corresponds to read addresses r0 and r1, and the number of read addresses is not greater than the number of address memories A0 and A1 corresponding to bank group 0; bank group 1 corresponds to read address r2, and the number of read addresses is not greater than the number of address memories A2 and A3 corresponding to bank group 1; the same applies to other bank groups. When the number of read addresses corresponding to each bank group is not greater than the number of corresponding address memories, the read addresses corresponding to each bank group can be written into the corresponding address memories within the same clock cycle. For example, within the same clock cycle, the read addresses r0 and r1 can be written into address memories A0 and A1 respectively, the read address r2 can be written into address memory A2, the read address r3 can be written into address memory A4, and the read address r3 can be written into address memory A6.
[0057] For example, in step S230, data can be read from the corresponding bankgroup according to the read addresses written in each address memory. For example, sub-data d0 corresponding to the read address r0 is read out from bank group 0, and sub-data d1 corresponding to the read address r1 is read out from bank group 0. The sub-data d0 and d1 can be used as a group of sub-data; sub-data d2 corresponding to the read address r2 is read out from bank group 1, and the sub-data d2 can be regarded as a group of sub-data; sub-data d3 corresponding to the read address r3 (not shown in the figure) is read out from bank group 2, and the sub-data d3 can be regarded as a group of sub-data; sub-data d4 corresponding to the read address r4 (not shown in the figure) is read out from bank group 3, and the sub-data d4 can be regarded as a group of sub-data. In this way, N groups of sub-data corresponding one-to-one to N groups of read addresses are obtained.
[0058] For example, in step S240, the N groups of sub-data corresponding to the same read request can be merged to obtain the target data corresponding to the read request. For example, the sub-data d0 to d4 can be merged, and the merged data can be used as the target data to be read out by the read request. Then, the target data is transmitted to the target device through a transmission device such as a bus. The target device can be the initiating device of the read request (also referred to as the requesting device), or other devices other than the initiating device. For example, it can be a main memory or a processor of a computing device. By merging the N groups of sub-data to obtain the target data corresponding to the read request, the bandwidth of the subsequent data bus can be aligned, the data is returned to the initiating device through one transmission, the number of transmissions of the transmission device is reduced, the utilization rate of the transmission device is improved, and the number of accesses of the initiating device is reduced, thereby improving the access bandwidth of the initiating device.
[0059] According to the data reading method of the embodiments of the present disclosure, by setting at least two address memories for each storage unit, multiple read addresses with read conflicts can be first written into the address memories of the corresponding storage units, reducing the backpressure on the data input end, solving the problem of read conflicts at the hardware level, and improving the parallel processing efficiency of the computing device. Moreover, by merging the return data corresponding to the read addresses with conflicts, the number of transmissions of the transmission device is reduced, and the utilization rate of the transmission device is improved.
[0060] For example, in step S230, reading sub-data corresponding to the read addresses in the N groups of address memories from N storage units to obtain N groups of sub-data corresponding to the N groups of read addresses respectively may include: in each clock cycle, based on the arbitration rule, selecting an address memory from the first group of address memories; reading the sub-data corresponding to the read address in the selected address memory from the first storage unit.
[0061] For example, the arbitration rule may adopt a polling method. In each clock cycle, based on the polling order of at least two address memories in the first group of address memories, one address memory is selected from at least two address memories, and the sub-data corresponding to the read address in this address memory is read.
[0062] For example, referring to Figure 3 , taking bank group 0 as an example, in one clock cycle, the sub-data d0 corresponding to the read address r0 stored in the address memory A0 can be read out from bank group 0. In the next clock cycle, the sub-data d1 corresponding to the read address r1 stored in the address memory A1 is read out from bank group 0. In the next cycle, the sub-data corresponding to the next read address stored in the address memory A0 is read out from bank group 0, and so on, cyclically reading the sub-data corresponding to the read addresses stored in the two address memories. Based on this method, the two read addresses that are parallelly stored can be read in adjacent clock cycles, avoiding the situation where the waiting time of some sub-data of the same read request is too long. For example, after reading all the sub-data corresponding to the read request, these sub-data can be merged.
[0063] For example, each address memory includes multiple storage locations, and the width of each storage location in the multiple storage locations corresponds to the data bit width of one read address. Among them, reading the sub-data corresponding to the read address in the selected address memory from the first storage unit includes: reading the sub-data corresponding to the first read address in the selected address memory from the first storage unit.
[0064] For example, as Figure 3 shown, taking the two address memories A0 and A1 corresponding to bank group 0 as an example, each address memory may include 4 storage locations, and the width of each storage location is equal to or greater than the data bit width of each read address.
[0065] For example, if the read address r0 and the read address r1 simultaneously correspond to bank group 0, the read addresses r0 and r1 are written into the address memories A0 and A1 in parallel. If there is no other read address in the address memory A0 or A1, the read address r0 or r1 is written into the first storage location of the address memory A0 or A1 respectively; if there is still an unwritten read address in the address memory A0 or A1, for example, there is another read address in the first storage location of the address memory A0, the read address r0 is written into the second storage location of the address memory A0.
[0066] For example, in each clock cycle, an address memory can be selected according to the polling order of address memories A0 and A1 and the policy of polling scheduling. After the address memory is selected, a read address can be selected according to the storage order of the read addresses in the selected address memory. Each time a read address is selected from the address memory, the read address stored in the first storage location can be selected. After the read address in the first storage location is read out, the read addresses in other storage locations can be uniformly moved forward.
[0067] For example, in some embodiments, the data reading method may further include: recording the correspondence between each read address in the read request and N groups of address memories to form record information.
[0068] For example, as Figure 3 shown, the address processing unit can be connected to the recording unit, and the address processing unit can record the address memory assigned to each read address in the recording unit. Therefore, by querying the information in the recording unit, it can be known which address memory each read address is assigned to.
[0069] For example, each storage unit corresponds to a group of data memories. Each group of data memories includes at least two data memories. The group of address memories and the group of data memories corresponding to each storage unit correspond to each other.
[0070] For example, as Figure 3 shown, for each bank group, a group of data memories is provided. The present disclosure embodiment does not limit the number of data memories included in the data address memory. Each group of data memories may include, for example, two or more data memories. Hereinafter, the case where each group of data memories includes two data memories will be taken as an example for description. For example, for bank group 0, data memories B0 to B1 are provided; for bank group 1, data memories B2 to B3 are provided; for bank group 2, data memories B4 to B5 are provided; for bank group 3, data memories B6 to B7 are provided,..., for bank group 7, data memories B14 to B15 are provided. For example, each data memory can be implemented as a device for storage such as a register. The number of address memories corresponding to each storage unit can be the same as the number of corresponding data memories, and the multiple address memories corresponding to each storage unit can correspond one-to-one to the corresponding multiple data memories.
[0071] Figure 4 The flowchart shows sub-data corresponding to reading out N groups of read addresses provided by at least one embodiment of the present disclosure.
[0072] As Figure 4As shown, in step S230, sub-data corresponding to the read addresses in the N address memories are read from the N storage units to obtain N groups of sub-data corresponding to the N groups of read addresses respectively, which may include steps S231 and S232.
[0073] Step S231: For each storage unit, after reading the sub-data corresponding to one read address each time, select one data memory from the corresponding group of data memories, and write the read sub-data into the selected data memory.
[0074] Step S232: Based on the recorded information, read the sub-data corresponding to the read request from the N groups of data memories to obtain N groups of sub-data.
[0075] For example, for each storage unit, after reading the sub-data corresponding to one read address each time, the sub-data can be stored in turn in the corresponding multiple data memories (where "multiple" means two or more) in a polling manner. In each clock cycle, based on the polling order of the multiple data memories, select one of them, and store the read sub-data in the data memory selected by polling. Taking bank group 0 as an example, bank group 0 corresponds to data memories B0 and B1. In one clock cycle, the sub-data d0 corresponding to the read address r0 is read. The read address r0 was previously stored in the address memory A0. Therefore, the sub-data d0 can be stored in the data memory B0 corresponding to the address memory A0. In the next clock cycle, the sub-data d1 corresponding to the read address r1 is read, and the sub-data d1 can be stored in the data memory B1. After reading the next sub-data, the sub-data is stored in the data memory B0, and so on, storing the read sub-data in the two data memories cyclically. In this way, the correspondence between the read address in the address memory and the sub-data in the data memory is also realized.
[0076] For example, each data memory includes multiple storage locations, and the width of each storage location among the multiple storage locations corresponds to the bit width of one sub-data. For example, as Figure 3 shown, taking the two data memories B0 and B1 corresponding to bank group 0 as an example, each data memory may include 4 storage locations, and the width of each storage location is equal to or greater than the bit width of each sub-data.
[0077] For example, after reading the sub-data d0 from bank group 0, if there is no other sub-data in the data memory B0, the sub-data d0 is written into the first storage location of the data memory B0; if there are still unwritten sub-data in the data memory B0, for example, there are other sub-data in the first storage location of the data memory B0, the sub-data d0 is written into the second storage location of the data memory B0.
[0078] For example, in each clock cycle, one data memory can be selected according to the polling order of data memories B0 and B1 and the strategy of polling scheduling. After the data memory is selected, the sub-data can be stored in the storage order of the read addresses in the selected data memory.
[0079] For example, the data processing unit can obtain sub-data from the recorded data memories according to the recorded information. Each time sub-data is obtained from a data memory, the sub-data stored at the first storage location can be selected. After the sub-data at the first storage location is read out, the sub-data at other storage locations can be uniformly moved forward.
[0080] For example, in step S232, based on the recorded information, reading out sub-data corresponding to the read request from N groups of data memories to obtain N groups of sub-data may include: when the number of addresses included in each of the N groups of read addresses is not greater than the number of memories included in the corresponding group of address memories, based on the recorded information, simultaneously reading out sub-data corresponding to the read request from N groups of data memories to obtain N groups of sub-data.
[0081] For example, when the number of addresses included in each of the N groups of read addresses is not greater than the number of memories included in the corresponding group of address memories, the N groups of read addresses can be written into the address memories in parallel, and the recording unit records which address memory each read address is stored in. After reading out the sub-data corresponding to each read address from the N storage units, the sub-data can be stored in the corresponding data memories. Therefore, the multiple sub-data corresponding to the multiple read addresses stored in parallel can also be stored at the same storage locations in the multiple data memories. Also, since there is a corresponding relationship between the address memory and the data memory, the sub-data corresponding to the read request can be determined to be stored in which data memories according to the recorded information. Based on this, multiple sub-data corresponding to the read request can be simultaneously read out from these data memories at one time.
[0082] For example, read addresses r0, r1, r2, r3, and r4 are written into address memories A0, A1, A2, A4, and A6 in parallel, and this information can be recorded in the recording unit. Sub-data d0, d1, d2, d3, and d4 corresponding to the read addresses r0, r1, r2, r3, and r4 are respectively written into data memories B0, B1, B2, B4, and B6. According to the information recorded in the recording unit, the data processing unit can gate data memories B0, B1, B2, B4, and B6 through a second selector (MUX) and simultaneously read the sub-data in these data memories, so as to obtain all the sub-data corresponding to the read request. After the data processing unit reads these sub-data, it can perform a merging operation on these sub-data to obtain the target data. Based on this method, the read sub-data can be stored in the expected data memories, and all the sub-data of the read request can be simultaneously read from these expected data memories by recording information, which improves the processing efficiency.
[0083] For example, in some other embodiments, in step S220, when the number of addresses in the first group of read addresses is greater than the number of memories in the first group of address memories corresponding to the first storage unit, the first group of read addresses is written into the first group of read addresses in multiple times. In step S232, when there is at least one group of read addresses among the N groups of read addresses whose number of included addresses is greater than the number of memories included in the corresponding group of address memories, the sub-data corresponding to the read request is read from the N groups of data memories in multiple times. In step S240, the sub-data read from the N groups of data memories each time is stored in the merging memory until all the sub-data corresponding to the read request are stored in the merging memory, and the sub-data in the merging memory is merged to obtain the target data corresponding to the read request.
[0084] Figure 5A and 5B shows another schematic diagram of the data reading process provided by at least one embodiment of the present disclosure.
[0085] Such as Figure 5A and 5BAs shown, taking bank group 0 as an example, bank group 0 corresponds to two address memories A0 and A1. There are more than two read addresses in the read request corresponding to bank group 0. For example, in addition to read addresses r0 and r1, there are also read addresses r5 and r6 corresponding to bank group 0. In this case, the read addresses r0, r1, r5, and r6 can be written into the address memories A0 and A1 in two times. For example, in one clock cycle, first write the read addresses r0 and r1 into the address memories A0 and A1 in parallel, and at the same time write the read addresses r2, r3, and r4 into the address memories A2, A4, and A6 respectively. In the next clock cycle, write the read addresses r5 and r6 into the address memories A0 and A1 in parallel. In the data reading stage, in a certain clock cycle, the sub-data d0, d1, d2, d3, and d4 are stored in the data memories B0, B1, B2, B4, and B6 respectively. The data processing unit can read out these sub-data according to the recorded information and store these sub-data in the merging memory in the data processing unit. After, for example, two clock cycles, the sub-data d0 and d1 corresponding to the read addresses r5 and r6 are stored in the data memories B0 and B1 respectively. The data processing unit can read out the sub-data d0 and d1 according to the recorded information and store the sub-data d0 and d1 in the merging memory. After the merging memory receives all the sub-data corresponding to the read request, these sub-data can be merged to obtain the target data. Based on this method, when the number of read addresses corresponding to a storage unit is more than the number of corresponding address memories, the reading and merging of sub-data can be simply and efficiently implemented by using the recorded information and N groups of data memories.
[0086] For example, in some other embodiments, when the number of addresses in the first set of read addresses is greater than the number of memories in the first set of address memories corresponding to the first storage unit, in addition to writing the first set of read addresses into the address memories in multiple batches and reading and merging the data in multiple batches based on the recorded information, other methods can also be adopted for processing. For example, if the number of addresses in the first set of read addresses is greater than the number of memories in the first set of address memories corresponding to the first storage unit, multiple read addresses in the first set of read addresses can be reallocated, and the extra read addresses can be reallocated to other storage units, so that multiple read addresses can be written into the address memories in parallel. For example, the first set of address memories corresponding to the first storage unit includes 2 address memories, and the first set of read addresses includes 3 read addresses. Two of the read addresses can be written into the 2 address memories in parallel, and the other read address can be reallocated to other idle storage units (storage units with idle address memories in the current cycle). For example, when setting the number of address memories corresponding to each storage unit, the number of read addresses that each storage unit usually corresponds to in one cycle is considered. Therefore, usually, the number of address memories corresponding to each storage unit can meet the storage requirements of the corresponding read addresses. When the number of read addresses is greater than the number of address memories, the method of dividing into batches or reallocation can be adopted for processing, so that the data reading method of the embodiments of the present application can handle various situations.
[0087] Figure 6 FIG. 4 shows a schematic block diagram of a data reading device 600 provided by at least one embodiment of the present disclosure.
[0088] For example, as Figure 6 shown, the data reading device 600 includes an address processing unit 610, M groups of address memories 620, M reading units 630, and a data processing unit 640. These components are interconnected through a bus system and / or other forms of connection mechanisms (not shown). For example, these modules / units / devices can be implemented through hardware (such as circuits) modules, software modules, or any combination of the two. The same applies to the following embodiments and will not be repeated. For example, these units can be implemented through a processing unit with data processing capabilities and / or instruction execution capabilities and corresponding computer instructions. It should be noted that Figure 6 the components and structures of the data reading device 600 shown are exemplary, not restrictive. According to needs, the data reading device 600 can also have other components and structures.
[0089] The M groups of address memories 620 respectively correspond to M storage units, and each group of address memories includes at least two address memories.
[0090] The address processing unit 610 is connected to the M groups of address memories 620, and the address processing unit 610 is configured to: receive a read request, where the read request includes N groups of read addresses respectively corresponding to N of the M memory cells, and each of the N groups of read addresses includes one or more read addresses; write the N groups of read addresses into the N groups of address memories corresponding to the N memory cells respectively, where the N groups of read addresses include a first group of read addresses corresponding to a first memory cell among the N memory cells, and in the case where the first group of read addresses includes at least two read addresses and the number of addresses in the first group of read addresses is not greater than the number of memories in the first group of address memories corresponding to the first memory cell, the first group of read addresses is written into the first group of address memories in parallel.
[0091] The M reading units 630 respectively correspond to the M memory cells, and the M reading units are configured to read sub-data corresponding to the read addresses in the N groups of address memories from the N memory cells to obtain N groups of sub-data respectively corresponding to the N groups of read addresses.
[0092] The data processing unit 640 is configured to merge the N groups of sub-data to obtain target data corresponding to the read request, and transmit the target data to the target device through the data transmission unit.
[0093] For example, the M groups of address memories 620 can refer to Figure 3 or Figure 5A the multi-group address memories shown. The address processing unit 610 can refer to Figure 3 or Figure 5A the address processing unit shown. The M reading units 630 can include an arbiter arb connected to the address memories as referred to in Figure 3 or Figure 5A shown. The arbiter arb can select one address memory from at least two connected address memories, and then the reading unit can read the sub-data corresponding to the read address in the selected address memory from the memory cell. The data processing unit 640 can refer to Figure 3 or Figure 5A the data processing unit shown.
[0094] For example, the M reading units include a first reading unit corresponding to the first memory cell, and the first reading unit is further configured to: in each clock cycle, select one address memory from the first group of address memories based on an arbitration rule; read the sub-data corresponding to the read address in the selected address memory from the first memory cell.
[0095] For example, the data reading device further includes a recording unit, refer to Figure 3 or Figure 5AThe recording unit shown, which is configured to record the correspondence between each read address in the read request and the N groups of address memories to form recording information.
[0096] For example, the data reading device further includes M groups of data memories and a transmission unit. The M groups of data memories respectively correspond to M storage units. Each group of data memories includes at least two data memories. The transmission unit is connected between the M groups of data memories and the data processing unit. The M groups of data memories can refer to Figure 3 or Figure 5A the multi-group data memories shown, and the transmission unit can refer to Figure 3 or Figure 5A the data selector (MUX) connected to the multi-group data memories shown.
[0097] For example, the M reading units 630 can also include referring to Figure 3 or Figure 5A the arbiter arb connected to the data memories shown. Each reading unit can be further configured to: after reading the sub-data corresponding to one read address from the corresponding storage unit each time, select one data memory from the corresponding group of data memories and write the read sub-data into the selected data memory.
[0098] For example, the transmission unit is configured to read the sub-data corresponding to the read request from the N groups of data memories based on the recording information and send the sub-data to the data processing unit.
[0099] For example, the above-mentioned unit / module / device can be hardware, software, firmware, and any feasible combination thereof. For example, the above-mentioned unit / module / device can be a dedicated or general-purpose circuit, chip, or device, etc., or can also be a combination of a processor and a memory. Regarding the specific implementation forms of the above-mentioned various units, the embodiments of the present disclosure do not limit this.
[0100] For example, the above-mentioned unit / module / device can include code and programs stored in a memory; the processor can execute the code and programs to implement some or all of the functions of the above-mentioned unit / module / device. For example, the above-mentioned unit / module / device can be a dedicated hardware device used to implement some or all of the functions of the above-mentioned unit / module / device. For example, the above-mentioned unit / module / device can be a circuit board or a combination of multiple circuit boards for implementing the above-mentioned functions. In the embodiments of the present disclosure, the combination of the one circuit board or multiple circuit boards can include: (1) one or more processors; (2) one or more non-transitory memories connected to the processor; and (3) firmware stored in the memory that can be executed by the processor.
[0101] It should be noted that in the embodiments of the present disclosure, each unit of the data reading device 600 corresponds to each step of the aforementioned data reading method. For the specific functions of the data reading device 600, reference can be made to the relevant descriptions of the data reading method, which will not be elaborated here. Figure 7 The components and structures of the data reading device 600 shown are exemplary and not restrictive. According to needs, the data reading device 600 may further include other components and structures. The data reading device 600 may include more or fewer circuits or units, and the connection relationships between the various circuits or units are not limited and can be determined according to actual requirements. The specific composition manners of the various circuits or units are not limited and can be composed of analog devices according to circuit principles, or composed of digital chips, or in other applicable manners.
[0102] At least one embodiment of the present disclosure further provides an electronic device, which includes a processor and a memory, and the memory stores one or more computer program modules. The one or more computer program modules are configured to be executed by the processor to implement the above-mentioned data reading method.
[0103] Figure 7 It is a schematic block diagram of an electronic device provided by some embodiments of the present disclosure. As Figure 7 shown, the electronic device 700 includes a processor 710 and a memory 720. The memory 720 stores non-transitory computer-readable instructions (such as one or more computer program modules). The processor 710 is used to run the non-transitory computer-readable instructions, and when the non-transitory computer-readable instructions are run by the processor 710, one or more steps in the above-mentioned data reading method are executed. The memory 720 and the processor 710 may be interconnected through a bus system and / or other forms of connection mechanisms (not shown). For the specific implementation of each step of the data reading method and the related explanatory content, reference can be made to the embodiments of the above data reading method, and the repeated parts will not be elaborated here.
[0104] It should be noted that Figure 7 the components of the electronic device 700 shown are exemplary and not restrictive. According to actual application needs, the electronic device 700 may further have other components.
[0105] For example, the processor 710 and the memory 720 may communicate with each other directly or indirectly.
[0106] For example, the processor 710 and the memory 720 may communicate through a network. The network may include a wireless network, a wired network, and / or any combination of a wireless network and a wired network. The processor 710 and the memory 720 may also communicate with each other through a system bus, and the present disclosure does not limit this.
[0107] For example, the processor 710 and the memory 720 may be disposed on the server side (or in the cloud).
[0108] For example, the processor 710 may control other components in the electronic device 700 to perform desired functions. For example, the processor 710 may be a central processing unit (CPU), a graphics processing unit (GPU), or other forms of processing units with data processing capabilities and / or program execution capabilities. For example, the central processing unit (CPU) may be of the X86 or ARM architecture, etc. The processor 710 may be a general-purpose processor or a dedicated processor, and may control other components in the electronic device 700 to perform desired functions.
[0109] For example, the memory 720 may include any combination of one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, etc. One or more computer program modules may be stored on the computer-readable storage media, and the processor 710 may run one or more computer program modules to implement various functions of the electronic device 700. Various application programs and various data, as well as various data used and / or generated by the application programs, etc. may also be stored in the computer-readable storage media.
[0110] It should be noted that in the embodiments of the present disclosure, the specific functions and technical effects of the electronic device 700 may refer to the description of the data reading method in the foregoing text, and will not be elaborated herein.
[0111] Figure 8 A schematic block diagram of another electronic device provided for some embodiments of the present disclosure. The electronic device 800 is, for example, suitable for implementing the data reading method provided by the embodiments of the present disclosure. The electronic device 800 may be a terminal device, etc. It should be noted that Figure 8 The illustrated electronic device 800 is only an example, and it will not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0112] Such as Figure 8As shown, the electronic device 800 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 810, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 820 or a program loaded from a storage device 880 into a random access memory (RAM) 830. In the RAM 830, various programs and data required for the operation of the electronic device 800 are also stored. The processing device 810, the ROM 820, and the RAM 830 are connected to each other through a bus 840. An input / output (I / O) interface 850 is also connected to the bus 840.
[0113] Generally, the following devices may be connected to the I / O interface 850: an input device 860 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 870 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 880 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 890. The communication device 890 may allow the electronic device 800 to communicate with other electronic devices wirelessly or wiredly to exchange data. Although Figure 8 the electronic device 800 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices, and the electronic device 800 may alternatively implement or have more or fewer devices.
[0114] For example, according to an embodiment of the present disclosure, the above data reading method may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for executing the above data reading method. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device 890, or installed from the storage device 880, or installed from the ROM 820. When the computer program is executed by the processing device 810, the functions defined in the data reading method provided by the embodiments of the present disclosure may be implemented.
[0115] At least one embodiment of the present disclosure also provides a computer-readable storage medium, which stores non-temporary computer-readable instructions, and when the non-temporary computer-readable instructions are executed by a computer, the above data reading method may be implemented.
[0116] Figure 9 Schematic diagram of a storage medium provided for some embodiments of the present disclosure. As Figure 9 shown, the storage medium 900 stores non-temporary computer-readable instructions 910. For example, when the non-temporary computer-readable instructions 910 are executed by a computer, one or more steps in the data reading method described above are executed.
[0117] For example, the storage medium 900 can be applied to the above-mentioned electronic device 700. For example, the storage medium 900 can be Figure 7 the memory 720 in the illustrated electronic device 700. For example, the relevant description of the storage medium 900 can refer to Figure 7 the corresponding description of the memory 720 in the illustrated electronic device 700, which will not be elaborated here.
[0118] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features with similar functions disclosed in the present disclosure.
[0119] In addition, although the operations are depicted in a specific order, this should not be construed as requiring the operations to be performed in the specific order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments.
[0120] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms for implementing the claims.
[0121] Regarding the present disclosure, the following points need to be noted:
[0122] (1) The drawings of the embodiments of the present disclosure only relate to the structures involved in the embodiments of the present disclosure, and other structures can refer to the general design.
[0123] (2) Without conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other to obtain new embodiments.
[0124] The above is only the specific implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. The protection scope of the present disclosure shall be subject to the protection scope of the claims.
Claims
1. A data reading method for reading data in M storage units, each of which corresponds to a group of address memories, each group of address memories including at least two address memories, wherein: The method comprises: Receiving a read request, wherein the read request includes N groups of read addresses corresponding to N storage units in the M storage units, respectively, wherein each group of the N groups of read addresses includes one or more read addresses; writing the N groups of read addresses into N groups of address memories corresponding to the N storage units respectively, wherein the N groups of read addresses include a first group of read addresses corresponding to a first storage unit among the N storage units, and in a case where the first group of read addresses includes at least two read addresses and the number of addresses of the first group of read addresses is not greater than the number of memories of the first group of address memories corresponding to the first storage unit, the first group of read addresses are written into the first group of address memories in parallel; Reading out sub-data corresponding to the read addresses in the N groups of address memories from the N storage units to obtain N groups of sub-data corresponding to the N groups of read addresses respectively; Merging the N groups of sub-data to obtain target data corresponding to the read request, and transmitting the target data to a target device through a data transmission device; The sub-data corresponding to the read addresses in the N groups of address memories are read out from the N storage units to obtain N groups of sub-data corresponding to the N groups of read addresses respectively, including: selecting an address memory from the first group of address memories based on an arbitration rule in each clock cycle; reading out the sub-data corresponding to the read address in the selected address memory from the first storage unit; The first group of read addresses includes a first read address and a second read address written in parallel into the first group of address memory, and the sub-data corresponding to the first read address and the sub-data corresponding to the second read address are read out in adjacent clock cycles; Wherein, M is an integer greater than 1, and N is an integer greater than 1 and less than M.
2. The data reading method according to claim 1, wherein: The first group of read addresses includes k read addresses, where k is an integer greater than or equal to 2; Writing the N groups of read addresses into N groups of address memories corresponding to the N storage units respectively includes: When the number of memories in the first group of address memories is greater than or equal to k, the k read addresses are written in parallel into the k address memories in the first group of address memories.
3. The data reading method according to claim 1, wherein: Writing the N groups of read addresses into N groups of address memories corresponding to the N storage units respectively includes: In a case where the number of addresses included in each group of the N groups of read addresses is not greater than the number of memories included in a corresponding group of address memories, the N groups of read addresses are written into the N groups of address memories in parallel.
4. The data reading method according to claim 1, wherein: Each of the address memories comprises a plurality of storage locations, and a width of each storage location in the plurality of storage locations corresponds to a data bit width of one of the read addresses; Wherein, in each clock cycle, based on the arbitration rule, selecting an address memory from the first group of address memories comprises: selecting one of the at least two address memories based on the polling order of at least two address memories in the first group of address memories; Among them, reading out the sub-data corresponding to the read address in the selected address memory from the first storage unit includes: reading out the sub-data corresponding to the first read address in the selected address memory from the first storage unit.
5. The data reading method according to any one of claims 1 to 4, further comprising: The correspondence between each read address in the read request and the N groups of address memories is recorded to form record information.
6. The data reading method according to claim 5, wherein: Each of the storage units corresponds to a group of data storages, each group of data storages includes at least two data storages, and a group of address storages and a group of data storages of each of the storage units correspond to each other; Wherein, reading out the sub-data corresponding to the read addresses in the N groups of address memories from the N storage units to obtain N groups of sub-data respectively corresponding to the N groups of read addresses includes: For each of the storage units, after reading out a sub-data corresponding to a read address each time, a data storage is selected from a corresponding group of data storages, and the read sub-data is written into the selected data storage; Based on the record information, sub-data corresponding to the read request are read out from N groups of the data storage devices to obtain the N groups of sub-data.
7. The data reading method according to claim 6, wherein: Based on the record information, reading out the sub-data corresponding to the read request from the N groups of data storage to obtain the N groups of sub-data includes: When the number of addresses included in each of the N groups of read addresses is not greater than the number of memories included in a corresponding group of address memories, based on the record information, sub-data corresponding to the read request are simultaneously read out from the N groups of data memories to obtain the N groups of sub-data.
8. The data reading method according to claim 6, wherein: Writing the N groups of read addresses into N groups of address memories corresponding to the N storage units respectively includes: When the number of addresses of the first group of read addresses is greater than the number of memories of the first group of address memories corresponding to the first storage unit, writing the first group of read addresses into the first group of read addresses in multiple times; Wherein, based on the record information, reading out the sub-data corresponding to the read request from the N groups of data storage devices includes: In the case that at least one of the N groups of read addresses includes a larger number of addresses than a corresponding group of address memories, reading out sub-data corresponding to the read request from the N groups of data memories in multiple times; The step of merging the N groups of sub-data to obtain target data corresponding to the read request includes: The sub-data read out from the N groups of data storage devices each time are stored in a merge memory until all the sub-data corresponding to the read request are stored in the merge memory, and the sub-data in the merge memory are merged to obtain target data corresponding to the read request.
9. A data reading device, comprising: M groups of address memories corresponding to the M storage units, each group of address memories including at least two address memories; An address processing unit connected to the M groups of address memories; M read units, corresponding to the M storage units; and Data processing unit; Wherein, the address processing unit is configured as follows: Receiving a read request, wherein the read request includes N groups of read addresses corresponding to N storage units in the M storage units, respectively, wherein each group of the N groups of read addresses includes one or more read addresses; writing the N groups of read addresses into N groups of address memories corresponding to the N storage units respectively, wherein the N groups of read addresses include a first group of read addresses corresponding to a first storage unit among the N storage units, and in a case where the first group of read addresses includes at least two read addresses and the number of addresses of the first group of read addresses is not greater than the number of memories of the first group of address memories corresponding to the first storage unit, the first group of read addresses are written into the first group of address memories in parallel; The M reading units are configured to read out sub-data corresponding to the read addresses in the N groups of address memories from the N storage units to obtain N groups of sub-data corresponding to the N groups of read addresses respectively; The data processing unit is configured to merge the N groups of sub-data to obtain target data corresponding to the read request, and transmit the target data to the target device through the data transmission unit; The M reading units include a first reading unit corresponding to the first storage unit, and the first reading unit is further configured to: select an address memory from the first group of address memories based on an arbitration rule in each clock cycle; read out sub-data corresponding to the read address in the selected address memory from the first storage unit; The first group of read addresses includes a first read address and a second read address written in parallel into the first group of address memory, and the sub-data corresponding to the first read address and the sub-data corresponding to the second read address are read out in adjacent clock cycles; Wherein, M is an integer greater than 1, and N is an integer greater than 1 and less than M.
10. The data reading device according to claim 9, further comprising: The recording unit is configured to record the corresponding relationship between each read address in the read request and the N groups of address memories to form recording information.
11. The data reading device according to claim 10, further comprising: M groups of data storage devices, corresponding to the M storage units, each group of data storage devices including at least two data storage devices; as well as A transmission unit connected between the M groups of data storage devices and the data processing unit; Each of the reading units is further configured to: after reading a sub-data corresponding to a read address from the corresponding storage unit each time, select a data storage from a corresponding group of data storages, and write the read sub-data into the selected data storage; The transmission unit is configured to read out sub-data corresponding to the read request from N groups of the data storage devices based on the record information, and send the sub-data to the data processing unit.
12. An electronic device comprising: processor; a memory storing one or more computer program modules; The one or more computer program modules are configured to be executed by the processor to implement the data reading method according to any one of claims 1 to 8.
13. A computer-readable storage medium storing non-transitory computer-readable instructions, which, when executed by a computer, can implement the data reading method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Processor and arithmetic processing method
CN116126216A
Cross access system based on AXI interface
CN117171070A
Method and device for writing address data, electronic equipment and storage medium
CN118469795A