Data storage device, data access method and electronic equipment

By adding a register module to the data storage device and performing a prefetch operation, the data of the target conflicting access request is transferred to the register in advance, solving the access delay problem caused by storage body conflicts, and achieving more efficient data reading and system performance improvement.

CN120669923APending Publication Date: 2025-09-19海光信息技术(成都)有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510846166.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

In a storage system, when multiple access requests simultaneously apply to access the same storage bank, storage bank conflicts lead to increased access latency and decreased system performance. Existing solutions cannot effectively avoid storage bank conflicts.

Method used

A register module is added to the data storage device, and the data of the target conflicting access request is transferred to the register in advance through a prefetch operation, and runs on a separate hardware path so that it is executed in parallel with other conflicting access requests to avoid storage body conflicts.

Benefits of technology

Through hardware optimization, the delay and waiting time caused by storage conflicts are reduced, the data reading efficiency is improved, the storage conflicts are avoided, and the system performance is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120669923A_ABST
    Figure CN120669923A_ABST
Patent Text Reader

Abstract

The invention discloses a data storage device, a data access method and electronic equipment. The data storage device comprises an analysis module, a memory bank module and a register module which are coupled. The memory bank module comprises a plurality of memory banks, and the register module comprises a plurality of registers. The analysis module is configured to obtain a plurality of access requests of a plurality of threads applying for accessing the memory bank module; analyzing the plurality of access requests to obtain access addresses corresponding to the plurality of access requests; determining whether the plurality of access requests comprise a plurality of conflict access requests of which the access addresses point to the same target memory bank in the plurality of memory banks or not; and in response to a plurality of conflict access requests included in the plurality of access requests, determining a target conflict access request from the plurality of conflict access requests, executing a prefetching operation, and prefetching target data in a target address corresponding to an access address of the target conflict access request into a target register corresponding to the target memory bank in the register module. The data storage device can reduce access delay and avoid memory stack conflicts.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to a data storage device, a data access method, and an electronic device. Background Art

[0002] In a storage system, memory can be divided into multiple banks to support parallel access. Each bank includes multiple memory cells, and each memory cell includes multiple memory elements. A memory element is the smallest physical unit of data storage, and each memory element can store, for example, a single binary value "0" or "1." A memory cell is used to store a continuous set of data. The number of binary bits that each memory cell can store is typically an integer multiple of 8, such as 8 bits, 16 bits, 32 bits, or 64 bits.

[0003] Each memory cell within a memory bank has a unique storage address. When an access request is made to access data in the memory, the target memory cell in the target memory bank is found based on the mapping between the access address and the storage address. Because each memory bank can perform independent read and write operations, multiple access requests can be issued simultaneously to different memory banks. These access requests can be processed in parallel by multiple memory banks, reducing latency and improving memory access efficiency and performance.

[0004] However, when multiple access requests simultaneously request access to the same memory bank, the bank cannot process them simultaneously. This is called a bank conflict. For example, a bank conflict occurs when multiple access requests simultaneously request access to different memory addresses within the same memory bank. In this case, these access requests must be processed serially, increasing access latency and degrading system performance. Summary of the Invention

[0005] At least one embodiment of the present disclosure provides a data storage device, which includes a coupled parsing module, a storage body module and a register module, wherein the storage body module includes multiple storage bodies, and the register module includes multiple registers corresponding to the multiple storage bodies. The parsing module is configured to: obtain multiple access requests from multiple threads applying to access the storage body module; parse the multiple access requests to obtain access addresses corresponding to the multiple access requests; determine whether the multiple access requests include multiple conflicting access requests whose access addresses point to the same target storage body in the multiple storage bodies; in response to the multiple access requests including the multiple conflicting access requests, determine a target conflicting access request from the multiple conflicting access requests, and perform a prefetch operation on the target conflicting access request, wherein the prefetch operation includes prefetching target data in a target address corresponding to the access address of the target conflicting access request into a target register corresponding to the target storage body in the register module.

[0006] For example, in the data storage device provided in at least one embodiment of the present disclosure, the register module is configured to: in response to executing the prefetch operation, save the target data in the target address and record the target address through the target register, so as to respond to the target conflict access request, and the target thread corresponding to the target conflict access request reads the target data from the target register according to the target address.

[0007] For example, in the data storage device provided in at least one embodiment of the present disclosure, the register module is further configured to: in response to the target thread reading the target data in the target register, clear the target data and the recorded target address stored in the target register.

[0008] For example, in the data storage device provided in at least one embodiment of the present disclosure, the multiple conflicting access requests include a first conflicting access request corresponding to a first thread and a second conflicting access request corresponding to a second thread, the first conflicting access request is determined to be the target conflicting access request, and the first thread reads the target data from the target register and the second thread reads the first data from the target address of the target storage body in parallel.

[0009] For example, the data storage device provided by at least one embodiment of the present disclosure also includes a scheduling module and a marking module, wherein the scheduling module is configured to perform a scheduling operation on the multiple access requests; the marking module is configured to perform a marking operation on the target conflicting access request so that the target conflicting access request skips the scheduling operation of the scheduling module.

[0010] For example, the data storage device provided by at least one embodiment of the present disclosure further includes: an arbitration module configured to arbitrate the multiple access requests so that each storage bank responds to at most one access request at the same time.

[0011] For example, in the data storage device provided in at least one embodiment of the present disclosure, the arbitration module includes a first arbitration unit and a second arbitration unit, the first arbitration unit is configured to map the multiple access requests to the multiple storage bodies according to the scheduling result of the scheduling operation, and the second arbitration unit is configured to retrieve the first data corresponding to the multiple access requests from the multiple storage bodies.

[0012] For example, in the data storage device provided in at least one embodiment of the present disclosure, in response to executing the prefetch operation, the first arbitration unit is further configured to determine the target storage body in the storage body module to which the target conflicting access request is to be mapped, and the target storage body is configured to send the target data in the target address to the second arbitration unit, and the second arbitration unit is further configured to send the target data to the target register in the register module.

[0013] For example, the data storage device provided by at least one embodiment of the present disclosure also includes: a sorting module, connected to the second arbitration unit and the register module, configured to obtain the first data corresponding to the multiple access requests from the second arbitration unit and / or obtain the target data corresponding to the target conflict access request from the register module, sort the format of the first data and / or the target data, and send the sorted first data and / or the target data to the general register connected to the data storage device.

[0014] For example, in the data storage device provided in at least one embodiment of the present disclosure, N registers in the register module are configured to correspond to the same memory bank, where N is an integer greater than or equal to 1.

[0015] At least one embodiment of the present disclosure also provides a data access method, which includes: obtaining multiple access requests from multiple threads applying for access to a storage body module, wherein the storage body module includes multiple storage bodies; parsing the multiple access requests to obtain access addresses corresponding to the multiple access requests; determining whether the multiple access requests include multiple conflicting access requests whose access addresses point to the same target storage body in the multiple storage bodies; in response to the multiple access requests including the multiple conflicting access requests, determining a target conflicting access request from the multiple conflicting access requests, and performing a prefetch operation on the target conflicting access request, wherein the prefetch operation includes prefetching target data in a target address corresponding to the access address of the target conflicting access request into a target register corresponding to the target storage body in a register module, and the register module includes multiple registers corresponding to the multiple storage bodies.

[0016] For example, the data access method provided by at least one embodiment of the present disclosure also includes: in response to executing the prefetch operation, saving the target data in the target address and recording the target address through the target register, so as to respond to the target thread corresponding to the target conflict access request reading the target data from the target register according to the target address.

[0017] For example, the data access method provided by at least one embodiment of the present disclosure further includes: in response to the target thread completing reading the target data in the target register, clearing the target data and the recorded target address stored in the target register.

[0018] For example, in the data access method provided by at least one embodiment of the present disclosure, the multiple conflicting access requests include a first conflicting access request corresponding to a first thread and a second conflicting access request corresponding to a second thread, the first conflicting access request is determined to be the target conflicting access request, and the first thread reads the target data from the target register and the second thread reads the first data from the target address of the target storage body in parallel.

[0019] For example, the data access method provided by at least one embodiment of the present disclosure further includes: performing a scheduling operation on the multiple access requests; and performing a marking operation on the target conflicting access request so that the target conflicting access request skips the scheduling operation.

[0020] For example, the data access method provided by at least one embodiment of the present disclosure further includes: arbitrating the multiple access requests so that each memory bank responds to at most one access request at the same time.

[0021] For example, the data access method provided by at least one embodiment of the present disclosure further includes: mapping the multiple access requests to the multiple storage bodies according to the scheduling result of the scheduling operation; and taking out the first data corresponding to the multiple access requests from the multiple storage bodies.

[0022] At least one embodiment of the present disclosure further provides an electronic device, comprising the data storage device described in any of the above embodiments of the present disclosure.

[0023] For example, the electronic device provided by at least one embodiment of the present disclosure also includes a graphics processor, wherein the data storage device is arranged on a stream multiprocessor of the graphics processor as a shared memory, and a stream processor connected to the data storage device is also arranged on the stream multiprocessor, and the stream processor is configured to send the multiple access requests to the data storage device. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly introduced below. Obviously, the drawings in the following description only relate to some embodiments of the present disclosure, rather than limiting the present disclosure.

[0025] Figure 1 is a schematic diagram of the structure of an exemplary graphics processor;

[0026] Figure 2 is a schematic diagram of an exemplary memory;

[0027] Figure 3A A schematic diagram of multiple threads accessing memory simultaneously;

[0028] Figure 3B for Figure 3A Schematic diagram of the access time consumption of multiple threads;

[0029] Figure 4 A schematic diagram of a pipeline blockage caused by a memory bank conflict;

[0030] Figure 5 A schematic block diagram of a data storage device provided in at least one embodiment of the present disclosure;

[0031] Figure 6A A schematic diagram of an exemplary data storage device for multiple threads to access simultaneously;

[0032] Figure 6B for Figure 6A Schematic diagram of the access time consumption of multiple threads;

[0033] Figure 7 A schematic structural diagram of another data storage device provided by at least one embodiment of the present disclosure;

[0034] Figure 8A Another schematic diagram of multiple threads accessing memory simultaneously;

[0035] Figure 8B A schematic diagram of another exemplary data storage device for multiple threads to access simultaneously;

[0036] Figure 8C for Figure 8B Schematic diagram of the access time consumption of multiple threads;

[0037] Figure 9 A flowchart of a data access method provided by at least one embodiment of the present disclosure;

[0038] Figure 10 A schematic structural diagram of an electronic device provided in at least one embodiment of the present disclosure; and

[0039] Figure 11 A schematic structural diagram of another electronic device provided in at least one embodiment of the present disclosure. DETAILED DESCRIPTION

[0040] In order to make the purpose, technical solutions and advantages of the embodiments of the present disclosure more clear, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings of the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the described embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.

[0041] Unless otherwise defined, technical or scientific terms used in this disclosure should have the ordinary meanings understood by a person of ordinary skill in the art to which this disclosure belongs. The terms "first," "second," and similar terms used in this disclosure do not denote any order, quantity, or importance, but are simply used to distinguish different components. Terms such as "include" or "comprising" mean that the element or object preceding the term includes the elements or objects listed after the term, and their equivalents, without excluding other elements or objects. Terms such as "connected" or "connected" are not limited to physical or mechanical connections but may include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly. It should be understood that the various steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, method embodiments may include additional steps and / or omit steps shown. The scope of this disclosure is not limited in this respect.

[0042] The present disclosure is described below using several specific embodiments. To maintain clarity and conciseness in the following description of the embodiments of the present disclosure, detailed descriptions of known functions and components may be omitted. When any component of an embodiment of the present disclosure appears in more than one drawing, that component is represented by the same or similar reference numeral in each drawing.

[0043] In parallel computing, computing tasks are typically performed by multiple threads, which can simultaneously access multiple memory banks within the processor's storage device. Examples of processors or accelerators capable of performing parallel computing include, but are not limited to, graphics processing units (GPUs), general-purpose graphics processing units (GPGPUs), and tensor processing units (TPUs).

[0044] Figure 1 FIG. 1 is a schematic diagram of the structure of an exemplary graphics processor. Figure 1 As shown, the graphics processor includes multiple streaming multiprocessors (SMs), a second-level high-speed (L2) cache connected to the multiple streaming multiprocessors, and global memory. For example, each streaming multiprocessor includes multiple streaming processors (SPs), shared memory, a first-level high-speed (L1) cache, and constant memory.

[0045] In a graphics processor, multiple streaming multiprocessors (also known as compute units) can execute tasks in parallel to achieve high-performance computing. For example, before execution, multiple threads are distributed to multiple streaming multiprocessors in the form of thread blocks. Each thread block is then split into multiple thread warps, which are executed by multiple streaming processors (also known as compute cores). For example, each thread warp can include multiple threads, which can be 32 or another number, but this is not a limitation in the embodiments of the present disclosure.

[0046] For example, global memory is a large on-board memory typically used to store large amounts of data. Despite its large capacity, global memory has slow access speeds and low bandwidth. Furthermore, data must travel a long path from global memory to the streaming multiprocessor, resulting in high latency and making it unsuitable for frequent data access. L2 cache is a high-speed cache within the graphics processor. It has a large capacity and offers faster access speeds than global memory, thus reducing memory access latency.

[0047] Shared memory, located in the streaming multiprocessor, provides a shared memory area between threads for each SM's currently executing thread blocks in parallel, enabling inter-thread interaction. Shared memory offers low latency and a small footprint. It is a type of on-chip storage. Compared to global memory and L2 cache, shared memory has a smaller capacity. However, because it is closer to the streaming processor, shared memory is faster to access and has a higher bandwidth than global memory. Shared memory can store important data, reducing the kernel's demand for global memory. It can also be considered a programmable software cache; for example, data in shared memory can be explicitly manipulated.

[0048] Figure 2 Schematic diagram of an exemplary memory. For example, the following description is made by taking the memory as a shared memory. Figure 2 As shown, this shared memory is a pseudo-dual-port memory, divided into 32 equal-sized banks (bank 0, bank 1, ..., bank 31). These banks can be accessed simultaneously. For example, data can be read from or written to multiple banks simultaneously. An arbitration unit in the memory selects a bank, for example, by sending an access request to the corresponding bank based on the mapping between the access request address and the bank address.

[0049] However, when multiple threads attempt to access different addresses of the same memory bank within the same clock cycle, the hardware cannot process these requests in parallel (the arbitration unit cannot operate simultaneously), resulting in a memory bank conflict. After a memory bank conflict occurs, multiple read and write requests must be resubmitted to the memory bank.

[0050] Figure 3A A schematic diagram of multiple threads accessing memory simultaneously. Figure 3B for Figure 3A Schematic diagram of the access time consumption of corresponding multiple threads.

[0051] like Figure 3A As shown, threads 0, 1, 2 and 3 executed in parallel apply for access to the memory at the same time, wherein threads 0, 1 and 2 apply for access to memory bank 0 at the same time, and thread 3 applies for access to memory bank 1. Figure 3B As shown, because threads 0 and 3 each request access to different memory banks, they can execute in parallel. However, threads 0, 1, and 2 simultaneously request access to the same memory bank 0, resulting in a bank conflict and preventing parallel execution. In other words, because only one thread's access request can be sent to the same memory bank within the same clock cycle, memory bank 0 can only process thread 0's access request first, then thread 1's, and finally thread 2's. This bank conflict causes threads 0, 1, and 2 to switch from parallel reads to serial reads in the hardware, extending the total access time. For example, the serial execution of threads 0 and 2 starts at TS and ends at Te, a delay of two clock cycles (ΔT1) compared to the parallel execution of threads 0 and 2.

[0052] Memory bank conflicts not only increase access latency but also block the pipeline, severely impacting system performance.

[0053] Figure 4 This is a schematic diagram of a pipeline blockage caused by a memory bank conflict. Figure 4 As shown, an instruction pipeline comprises a five-stage pipeline, in which each instruction can be issued every clock cycle and executed within a fixed time (e.g., five clock cycles). The execution of each instruction is divided into five steps: the instruction fetch (IF) stage, the decode (ID) stage, the execute (EXE) stage, the memory access (MEM) stage, and the writeback (WB) stage. In the IF stage, a specified instruction is retrieved from the instruction cache. A portion of the retrieved instruction specifies the source register that can be used to execute the instruction. In the ID stage, the instruction is decoded and control logic is generated to retrieve the contents of the specified source register. Based on the control logic, operations are performed in the EXE stage using the retrieved contents. In the MEM stage, executing the instruction can read and write memory in the data cache. Finally, in the WB stage, the value obtained by executing the instruction can be written back to the general register.

[0054] For example, Figure 4 As shown, instructions W0, W1, W2 and W3 are instructions that need to be executed, and they will all go to the memory to access the storage body. When there is a storage body conflict when instruction W0 accesses the memory, the entire pipeline will be blocked, and the longer the blocking time, the deeper the impact on the performance of the pipeline.

[0055] In order to avoid memory conflicts, the inventors of the present disclosure have noticed that one solution is to increase the safe distance of instruction issuance through software dynamic scheduling, by selectively scheduling some non-memory access instructions so that they are interspersed between memory access instructions, thereby separating the conflicting instructions, and using the idle cycles between instructions to hide memory conflicts, thereby reducing the probability of memory conflicts. Another solution is to evenly distribute data in different memory banks by filling blank data or changing the address mapping relationship, so that addresses that originally pointed to the same memory bank are mapped to different memory banks, thereby avoiding memory conflicts. However, these solutions do not change the memory access pattern of the thread within the instruction. When faced with a large number of memory access conflicting instructions, it is impossible to directly avoid memory conflicts through hardware optimization.

[0056] At least one embodiment of the present disclosure provides a data storage device, a data access method and an electronic device. By adding a register module in the data storage device, when faced with multiple conflicting access requests for access to the same storage body, the data in the target address of the storage body corresponding to the target conflicting access request is transferred to the register in advance, and the target conflicting access request is transferred to a separate hardware path for operation, so that it does not access the storage body module and directly reads data from the target register in the register module, so that the target conflicting access request and other conflicting access requests can be executed in parallel, avoiding storage body conflicts, improving data reading efficiency, and reducing delays and waiting time.

[0057] Figure 5 A schematic block diagram of a data storage device provided in at least one embodiment of the present disclosure. Figure 5 As shown, at least one embodiment of the present disclosure provides a data storage device 500 including a parsing module 510, a storage module 520, and a register module 530. For example, these modules 510-530 are coupled to each other and can be implemented at least partially by hardware, firmware, and / or software. For example, when implemented by hardware, they can be implemented by digital circuits and / or analog circuits.

[0058] For example, the memory module 520 includes multiple memory banks. The number of memory banks in the memory module 520 can be divided into several, dozens, hundreds, thousands, or more according to actual needs. The embodiments of the present disclosure do not limit the number of memory banks in the memory module. For example, the number of memory banks in the memory module can be changed, for example, the number of memory banks in the memory module can be managed by software.

[0059] For example, the register module 530 includes multiple registers corresponding to multiple memory banks. For example, N registers can be configured for each memory bank, where N is an integer greater than or equal to 1. In other words, one or more registers can be configured for each memory bank. For example, the storage capacity of each register can be the same as the storage capacity of one or more storage units in the memory bank, or the storage capacity of each register can be different from the storage capacity of one or more storage units in the memory bank. The storage capacity of the registers can be set according to actual needs.

[0060] For example, the parsing module 510 is configured to obtain multiple access requests from multiple threads for accessing a storage module and parse the multiple access requests to obtain access addresses corresponding to the multiple access requests.

[0061] In an embodiment of the present disclosure, in addition to the parsing function, the parsing module 510 is further configured to determine whether multiple access requests include multiple conflicting access requests whose access addresses point to the same target storage bank among multiple storage banks. For example, the access address may be a logical address, and the storage bank to which the access address points may be determined based on a mapping rule. For example, one mapping rule may be to determine the storage bank number corresponding to the access address based on an index of the access address (e.g., X consecutive bits, where X is greater than or equal to 1). The embodiments of the present disclosure do not limit the specific mapping rule.

[0062] For example, in response to the multiple access requests including multiple conflicting access requests, the parsing module 510 determines a target conflicting access request from the multiple conflicting access requests and performs a prefetch operation on the target conflicting access request. For example, the prefetch operation includes prefetching target data at a target address corresponding to an access address of the target conflicting access request into a target register corresponding to a target memory bank in the register module.

[0063] Figure 6A A schematic diagram of an exemplary data storage device for simultaneous access by multiple threads. For example, the storage module 520 includes four storages, namely storage 0, storage 1, storage 2, and storage 3. When multiple threads executed in parallel simultaneously apply to access the storage module 520, for example, these threads are thread 0, thread 1, thread 2, and thread 3, the parsing module 510, after parsing the access addresses of the access requests of these threads, determines whether these access requests include conflicting access requests. For example, if the access request 0 of thread 0, the access address of the access request 1 of thread 1, and the access address of the access request 2 of thread 2 all point to storage 0, and the access address of the access request 3 of thread 3 points to storage 1, then the parsing module 510 can determine that the access request 0 of thread 0, the access request 1 of thread 1, and the access request 2 of thread 2 are conflicting access requests.

[0064] Then, the parsing module 510 selects one as the target conflicting access request from access request 0, access request 1, and access request 2. For example, the parsing module 510 may determine access request 0 as the target conflicting access request according to the order of threads, and then perform a prefetch operation on access request 0.

[0065] For example, the prefetch operation can pre-fetch the target data at the target address to which the access address of access request 0 was originally mapped, and store the target data to be read by access request 0 in the register, so that access request 0 of thread 0 can directly read the target data from the register. Access request 1 of thread 1 can read data from memory bank 0, and access request 2 of thread 2 needs to wait until thread 1 completes the access operation before executing.

[0066] like Figure 6B As shown in FIG, through the hardware optimization of the data storage device provided by the embodiment of the present disclosure, thread 0 and thread 1 can be executed in parallel, that is, thread 0 reads target data from the register, and thread 1 reads data from memory bank 0, so that the delay is reduced from two clock cycles ( Figure 3B The ΔT1 in the clock is reduced to one clock period (ΔT2).

[0067] The number of registers in the register module can be set according to actual needs. For example, multiple registers can be configured for a memory bank. For example, in another example, two registers can be configured for memory bank 0, then the parsing module 510 can determine the access requests of thread 0 and thread 1 as target conflicting access requests, and perform prefetch operations on both access request 0 and access request 1. For example, the target data to be read by access request 0 and access request 1 are both taken out in advance from the corresponding target address in memory bank 0, and stored in the two registers corresponding to memory bank 0, so that access request 0 of thread 0 and access request 1 of thread 1 can directly read the target data from these two registers, and access request 2 of thread 2 can read data from memory bank 0, thereby realizing parallel reading of thread 0, thread 1 and thread 2, which can completely avoid memory bank conflicts, eliminate access delays, and improve data reading efficiency.

[0068] For example, in at least one embodiment of the present disclosure, the register module 530 is configured to, in response to executing a prefetch operation, save the target data in the target address and record the target address through the target register, so as to respond to the target thread corresponding to the target conflicting access request reading the target data from the target register according to the target address. For example, in the above example, the two registers are register a and register b, register a records the target address S00 in the memory bank 0 corresponding to access request 0 and saves the target data in the target address S00, and register b records the target address S01 in the memory bank 0 corresponding to access request 1 and saves the target data in the target address S01, so that when access request 0 and access request 1 are sent to the register module 530 at the same time, register a can respond to the read operation of thread 0 corresponding to access request 0, and register b can respond to the read operation of thread 1 corresponding to access request 1.

[0069] For example, in at least one embodiment of the present disclosure, the register module 530 is further configured to clear the target data stored in the target register and the recorded target address in response to the target thread completing reading the target data in the target register. For example, when the target register responds to the first prefetch operation, because the access address of the target thread (e.g., thread 0) corresponds to the address of the mth row in the target memory bank (e.g., memory bank 0), the target register will record the address of the mth row in memory bank 0 and the target data dm stored in the mth row address. After thread 0 completes reading the target data dm from the target register, the target register may clear the target data dm and the recorded address of the mth row, and then wait for a response to the second prefetch operation. For example, in response to the second prefetch operation, the access address of the target thread (for example, thread 4) may correspond to the nth row address of memory bank 0. Then the target register will record the nth row address of memory bank 0 and the target data dn stored in the nth row address. After thread 4 finishes reading the target data dn from the target register, the target register can clear the target data dn and the recorded nth row address, and then wait to respond to the next prefetch operation.

[0070] Figure 7 This is a structural diagram of another data storage device provided by at least one embodiment of the present disclosure. Figure 7 As shown, at least one embodiment of the present disclosure provides a data storage device 700, which includes a coupled parsing module 710, a memory bank module 720, a register module 730, a marking module 740, and a scheduling module 750. For example, the memory bank module 720 includes a plurality of memory banks 0-4, and the register module 730 includes a plurality of registers 0-4 corresponding one-to-one to the plurality of memory banks 0-4.

[0071] For example, reference Figure 7In S1, the parsing module 710 is configured to obtain multiple access requests from multiple threads applying for access to the storage body module 720 from the stream processor (SP) and parse the multiple access requests to obtain access addresses corresponding to the multiple access requests. In an embodiment of the present disclosure, in addition to the parsing function, the parsing module 710 is also configured with a function of determining whether to perform a prefetch operation. For example, the parsing module 710 can determine whether the multiple access requests include multiple conflicting access requests whose access addresses point to the same target storage body in multiple storage bodies; in response to the multiple access requests including multiple conflicting access requests, a target conflicting access request is determined from the multiple conflicting access requests, and a prefetch operation is performed on the target conflicting access request, wherein the prefetch operation includes prefetching the target data in the target address corresponding to the access address of the target conflicting access request to the target register of the corresponding target storage body in the register module.

[0072] about Figure 7 The detailed description of the parsing module 710, the storage module 720 and the register module 730 can be referred to in the above Figure 5 The detailed description of the parsing module 510, the storage module 520 and the register module 530 in the illustrated embodiment will not be repeated here.

[0073] For example, the scheduling module 750 is configured to perform a scheduling operation on multiple access requests. For example, the marking module 740 is configured to perform a marking operation on target conflicting access requests so that the target conflicting access requests skip the scheduling operation of the scheduling module 750 .

[0074] For example, reference Figure 7 In S2-1, the parsing module 710 sends multiple access requests to the marking module 740, the marking module 740 marks the access requests that are determined to be target conflicting access requests, and then sends these access requests to the scheduling module 750, as shown in FIG. Figure 7 As shown in S3-1 in FIG. After receiving multiple access requests, the scheduling module 750 skips checking the access requests with tags, that is, it only selects the access requests without tags from the multiple access requests for scheduling. For example, these access requests without tags may include conflicting access requests without tags.

[0075] For example, when the access requests include multiple conflicting access requests, the scheduling module 750 may adjust the conflicting access requests to serial reads, such as Figure 3BFor example, the scheduling module 750 can also adjust the processing time slots between threads through software scheduling methods. For example, the scheduling module 750 can read instructions in the program, determine whether there are dependencies between the instructions, and insert unrelated instructions between multiple instructions with memory bank conflicts, thereby disrupting the execution order of the instructions and ensuring that each memory access instruction has sufficient time to complete the use and clearing of the register.

[0076] For example, in one example, access requests 0-2 from threads 0-2 all access memory bank 0, and are therefore identified by parsing module 710 as conflicting access requests. Parsing module 710 selects a target conflicting access request from access requests 0-2 based on the number of registers corresponding to memory bank 0. If there is only one register corresponding to memory bank 0, parsing module 710 selects only one access request (e.g., access request 0) from access requests 0-2 as the target conflicting access request, and marking module 740 marks only this target conflicting access request (access request 0). After these access requests 0-2 are sent to scheduling module 750, scheduling module 750 filters out untagged access requests 1 and 2 from these access requests 0-2. Since access requests 1 and 2 are conflicting access requests to memory bank 0, scheduling module 750 still needs to schedule access requests 1 and 2. For example, the scheduling result is to first send access request 1 to memory bank 0, and then send access request 2 to memory bank 0 after the read operation of access request 1 completes.

[0077] For example, in some embodiments of the present disclosure, Figure 7 As shown, the data storage device 700 further includes an arbitration module configured to arbitrate multiple access requests so that each memory bank responds to at most one access request at a time.

[0078] For example, Figure 7 As shown, the arbitration module of the data storage device 700 includes a first arbitration unit 761 and a second arbitration unit 762. For example, the first arbitration unit 761 is configured to map multiple access requests to multiple storage banks based on the scheduling result of the scheduling operation, and the second arbitration unit 762 is configured to retrieve the first data corresponding to the multiple access requests from the multiple storage banks.

[0079] For example, in the above example, the scheduling module 750 skips the scheduling of access request 0 and determines that access request 1 is executed before access request 2. Then, the scheduling module 750 will send access request 1 to the first arbitration unit 761 first. Figure 7As shown in S4-1 in FIG. First arbitration unit 761 maps access request 1 to address 1 in memory bank 0, corresponding to access address 1, based on access address 1 of access request 1 and a preset mapping rule. Address 1 in memory bank 0 then sends the first data to be read by access request 1 to second arbitration unit 762. Then, after thread 1 finishes reading data from memory bank 0, scheduling module 750 sends access request 2 to first arbitration unit 761. First arbitration unit 761, based on access address 2 of access request 2 and a preset mapping rule, maps access request 2 to address 2 in memory bank 0, corresponding to access address 2.

[0080] For example, in some embodiments of the present disclosure, in response to executing a prefetch operation, the first arbitration unit 761 is further configured to determine a target memory bank in the memory bank module 720 to which the target conflicting access request is to be mapped, and the target memory bank is configured to send the target data in the target address to the second arbitration unit 762, and the second arbitration unit 762 is further configured to send the target data to the target register in the register module 730, such as Figure 7 As shown in S3-2 in .

[0081] For example, in the above example, the parsing module 710 determines that the access request 0 is a target conflict access request and needs to perform a prefetch operation on the access request 0. Then, the first arbitration unit 761 will determine in advance the target address in the memory bank 0 to which the access request 0 is to be mapped according to the preset mapping rule. Then, the memory bank 0 sends the target data in the target address to the second arbitration unit 762. The second arbitration unit 762 sends the target data to the target register corresponding to the memory bank 0 in the register module 720. The target register stores the target data and records the target address in the memory bank 0 to wait for the read operation of the thread 0 corresponding to the access request 0. For example, the access request 0 can skip the scheduling operation of the scheduling module 750 and directly read the target data from the target register of the register module 730, as shown in FIG. Figure 7 As shown in S4-2.

[0082] For example, in some embodiments of the present disclosure, Figure 7 As shown, the data storage device 700 further includes a sorting module 770. The sorting module 770 is connected to the second arbitration unit 762 and the register module 730, and is configured to obtain the first data corresponding to the plurality of access requests from the second arbitration unit 762, such as Figure 7 As shown in S5-1 in FIG, and / or obtaining target data corresponding to the target conflict access request from the register module 730, such as Figure 7Then, the sorting module 770 can sort the format of the first data and / or target data, and send the sorted first data and / or target data to the general register GPR connected to the data storage device 700, as shown in S5-2. Figure 7 As shown in S7.

[0083] Figure 8A The diagram is another diagram of multiple threads accessing memory simultaneously. For example, the memory module 720 includes four memory banks, namely, memory bank 0, memory bank 1, memory bank 2, and memory bank 3. For example, when multiple threads executed in parallel simultaneously apply to access the memory module 720, such as thread 0, thread 1, thread 2, and thread 3, the parsing module 710, after parsing the access addresses of the access requests of these threads, determines that the access addresses of thread 0's access request 0 and thread 1's access request 1 in these access requests both point to memory bank 0, and the access addresses of thread 2's access request 2 and thread 3's access request 3 both point to memory bank 1. Therefore, thread 0's access request 0 and thread 1's access request 1 are conflicting access requests for accessing memory bank 0, and thread 2's access request 2 and thread 3's access request 3 are conflicting access requests for accessing memory bank 1.

[0084] For example, the parsing module 710 selects access request 0 from access request 0 and access request 1 as the target conflicting access request, selects access request 2 from access request 2 and access request 3 as the target conflicting access request, and then performs a prefetch operation on access request 0 and access request 2, such as Figure 7 As shown in S2-2 in .

[0085] For example, the first arbitration unit 761 determines the target address 0 in the memory bank 0 to which access request 0 is to be mapped, and the target address 2 in the memory bank 1 to which access request 2 is to be mapped. Then, the memory bank 0 in the memory bank module 720 sends the target data 0 in the target address 0 to the second arbitration unit 762, and the second arbitration unit 762 pre-fetches the target data 0 into the register 0 corresponding to the memory bank 0. Similarly, the memory bank 1 sends the target data 2 in the target address 2 to the second arbitration unit 762, and the second arbitration unit 762 pre-fetches the target data 2 into the register 1 corresponding to the memory bank 1. Figure 7 As shown in S3-2 in .

[0086] For example, after the register module 730 receives target data 0 and target data 2 from the second arbitration unit 762, it instructs register 0 to save target data 0 and record target address 0, instructs register 1 to save target data 2 and record target address 2, and waits for the read operations of thread 0 and thread 2 corresponding to access request 0 and access request 2.

[0087] For example, after determining that access request 0 and access request 2 are target conflicting access requests, the marking module 740 marks access request 0 and access request 2, and sends multiple access requests 0 to 4 to the scheduling module 750, such as Figure 7 As shown in S3-1 in .

[0088] For example, after the scheduling module 750 identifies the tags in the multiple access requests, it sends the access request 1 and the access request 3 to the storage module 720 through the first arbitration unit 761. Figure 7 As shown in S4-2 in FIG, access request 0 and access request 2 are sent to register module 720, as shown in FIG. Figure 7 As shown in S4-2.

[0089] For example, the memory bank 0 in the memory bank module 720 sends the first data 1 corresponding to the access request 1 to the second arbitration unit 762, and the memory bank 1 sends the first data 3 corresponding to the access request 3 to the second arbitration unit 762. The second arbitration unit 762 sends the first data corresponding to the access request 1 and the access request 3 to the sorting module 770. Figure 7 At the same time, register 0 in register module 730 sends target data 0 to sorting module 770, and register 1 sends target data 2 to sorting module 770, as shown in S5-1 in FIG. Figure 7 As shown in S5-2.

[0090] For example, the sorting module 770 sorts the target data 0, first data 1, target data 2 and first data 3 corresponding to the access requests 0 to 4 of threads 0 to 4, and sends them to the general register GPR for subsequent processing, such as Figure 7 As shown in S7.

[0091] In summary, the data storage device provided by at least one embodiment of the present disclosure introduces a prefetch path (such as Figure 7 (as indicated by the double-line arrows in the figure), can respond to thread read operations in parallel with the memory module, thereby alleviating memory conflicts. Furthermore, the data storage device provided by at least one embodiment of the present disclosure is hardware-optimized and can be used in conjunction with other solutions (such as the software dynamic scheduling or blank data filling solutions mentioned above), thus providing strong compatibility.

[0092] Figure 8B FIG. 1 is a schematic diagram of another exemplary data storage device accessed by multiple threads simultaneously. Figure 8BAs shown, after the hardware optimization of the data storage device 700, thread 0 can directly read target data 0 from register 0 corresponding to memory bank 0, thereby avoiding the memory bank conflict caused by accessing memory bank 0 at the same time as thread 1; similarly, thread 2 can directly read target data 2 from register 1 corresponding to memory bank 1, thereby avoiding the memory bank conflict caused by accessing memory bank 1 at the same time as thread 3.

[0093] Figure 8C for Figure 8B The corresponding schematic diagram of the access time consumption of multiple threads. Figure 8C As shown, after the hardware optimization of the data storage device 700, parallel reading of thread 0, thread 1, thread 2 and thread 3 can be achieved, which can completely avoid storage bank conflicts, eliminate access delays, and improve data reading efficiency.

[0094] At least one embodiment of the present disclosure further provides a data access method for the above-mentioned data storage device 500 or 700. Figure 9 Flowchart of a data access method provided by at least one embodiment of the present disclosure. Figure 9 As shown, the data access method includes the following steps S100 to S400.

[0095] Step S100: obtaining multiple access requests from multiple threads for accessing a storage module, wherein the storage module includes multiple storages;

[0096] Step S200: parsing multiple access requests to obtain access addresses corresponding to the multiple access requests;

[0097] Step S300: determining whether the multiple access requests include multiple conflicting access requests whose access addresses point to the same target storage bank among the multiple storage banks;

[0098] Step S400: In response to multiple access requests including multiple conflicting access requests, a target conflicting access request is determined from the multiple conflicting access requests, and a prefetch operation is performed on the target conflicting access request, wherein the prefetch operation includes prefetching target data in a target address corresponding to an access address of the target conflicting access request into a target register corresponding to a target storage body in a register module, and the register module includes multiple registers corresponding to the multiple storage bodies.

[0099] For example, in at least one example of the embodiments of the present disclosure, the data access method also includes: in response to performing a prefetch operation, saving the target data in the target address and recording the target address through the target register, so as to respond to the target thread corresponding to the target conflicting access request reading the target data from the target register according to the target address.

[0100] For example, in at least one example of the embodiments of the present disclosure, the data access method further includes: in response to the target thread completing reading the target data in the target register, clearing the target data and the recorded target address stored in the target register.

[0101] For example, in at least one example of the embodiments of the present disclosure, multiple conflicting access requests include a first conflicting access request corresponding to a first thread and a second conflicting access request corresponding to a second thread, the first conflicting access request is determined to be a target conflicting access request, and the first thread reads the target data from the target register and the second thread reads the first data from the target address of the target storage body in parallel.

[0102] For example, in at least one example of the embodiments of the present disclosure, the data access method further includes: performing a scheduling operation on multiple access requests; and performing a marking operation on the target conflicting access request so that the target conflicting access request skips the scheduling operation.

[0103] For example, in at least one example of the embodiments of the present disclosure, the data access method further includes: arbitrating multiple access requests so that each memory bank responds to at most one access request at the same time.

[0104] For example, in at least one example of the embodiments of the present disclosure, the data access method further includes: mapping multiple access requests to multiple storage bodies according to the scheduling result of the scheduling operation; and taking out the first data corresponding to the multiple access requests from the multiple storage bodies.

[0105] It should be noted that the parsing module, marking module, scheduling module and sorting module in the embodiments of the present disclosure can be implemented by hardware, software, firmware and any feasible combination thereof, and the embodiments of the present disclosure do not limit this.

[0106] The data access method provided by at least one embodiment of the present disclosure can reduce memory bank conflicts, improve data reading efficiency, and reduce delays and waiting time.

[0107] At least one embodiment of the present disclosure further provides an electronic device, which includes the above-mentioned data storage device 500 or data storage device 700.

[0108] Figure 10 This is a schematic diagram of the structure of an electronic device provided by at least one embodiment of the present disclosure. Figure 10As shown, electronic device 1000 includes a graphics processing unit (GPU) 1001, which includes multiple streaming multiprocessors (SMs) 1002. Each streaming multiprocessor 1002 may include multiple stream processors (SPs) 1003, and a shared memory 1004 connected to the multiple stream processors 1003. For example, data storage device 500 or 700 is provided on the streaming multiprocessor 1002 of GPU 1001 to serve as shared memory 1004. Stream processor 1003 is configured to send multiple access requests to shared memory 1004.

[0109] The technical effect of the electronic device provided by at least one embodiment of the present disclosure is the same as the technical effect of the above-mentioned data storage device, and will not be repeated here.

[0110] Some embodiments of the present disclosure further provide an electronic device, which includes the data storage device of any of the above embodiments or can execute the data access method of any of the above embodiments.

[0111] Figure 11 A schematic diagram of the structure of another electronic device provided for at least one embodiment of the present disclosure. The electronic devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 11 The electronic device 1100 shown is merely an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0112] For example, Figure 11 As shown, in some examples, electronic device 1100 includes a processing device (e.g., a central processing unit, a graphics processor, etc.) 1101. This processing device 1101 may include a graphics processor according to any of the aforementioned embodiments, which may include a data storage device according to any of the aforementioned embodiments. Processing device 1101 can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) 1102 or programs loaded from storage device 1108 into random access memory (RAM) 1103. RAM 1103 also stores various programs and data required for the operation of the computer system. Processor 1101, ROM 1102, and RAM 1103 are connected to each other via a bus 1104. An input / output (I / O) interface 1105 is also connected to bus 1104.

[0113] For example, the following components may be connected to the I / O interface 1105: an input device 1106 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1107 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1108 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1109 which may also include, for example, a network interface card such as a LAN card, a modem, etc. The communication device 1109 may allow the electronic device 1100 to communicate with other devices wirelessly or by wire to exchange data, performing communication processing via a network such as the Internet. A drive 1110 is also connected to the I / O interface 1105 as needed. A removable storage medium 1111, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1110 as needed, so that a computer program read therefrom is installed into the storage device 1108 as needed. Although Figure 11 The electronic device 1100 is shown as including various devices, but it should be understood that it is not required to implement or include all of the devices shown, and more or fewer devices may be implemented or included instead.

[0114] For example, the electronic device 1100 may further include a peripheral interface (not shown in the figure), etc. The peripheral interface may be various types of interfaces, such as a USB interface, a lightning interface, etc. The communication device 1109 may communicate with a network and other devices through wireless communication, such as the Internet, an intranet, and / or a wireless network such as a cellular telephone network, a wireless local area network (LAN), and / or a metropolitan area network (MAN). Wireless communications may use any of a variety of communication standards, protocols, and technologies, including, but not limited to, Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (W-CDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Wi-Fi (e.g., based on IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, and / or IEEE 802.11n standards), Voice over Internet Protocol (VoIP), Wi-MAX, protocols for email, instant messaging, and / or Short Message Service (SMS), or any other suitable communication protocol.

[0115] For example, the electronic device 1100 can be any device such as a mobile phone, tablet computer, laptop computer, e-book, game console, television, digital photo frame, navigator, etc., or it can be any combination of data processing devices and hardware, and the embodiments of the present disclosure are not limited to this.

[0116] Although the present disclosure has been described in detail above using general descriptions and specific embodiments, it will be apparent to those skilled in the art that modifications or improvements may be made based on the embodiments of the present disclosure. Therefore, such modifications or improvements, as long as they do not depart from the spirit of the present disclosure, are within the scope of protection claimed by the present disclosure.

[0117] In addition to the above exemplary contents, the following points need to be explained in this disclosure:

[0118] (1) The drawings of the embodiments of the present disclosure only relate to the structures related to the embodiments of the present disclosure. Other structures may refer to conventional designs.

[0119] (2) For the sake of clarity, the thickness of layers or regions in the drawings used to describe the embodiments of the present disclosure are enlarged or reduced, that is, these drawings are not drawn according to the actual scale.

[0120] (3) In the absence of conflict, the embodiments of the present disclosure and the features therein may be combined with each other to form new embodiments.

[0121] The above description is only a specific embodiment of the present disclosure, but the protection scope of the present disclosure is not limited thereto. The protection scope of the present disclosure shall be based on the protection scope of the claims.

Claims

1. A data storage device, comprising a coupled parsing module, a storage module, and a register module, wherein: The memory module includes a plurality of memory banks, and the register module includes a plurality of registers corresponding to the plurality of memory banks. The parsing module is configured as follows: Acquire multiple access requests from multiple threads to access the storage module; Parsing the multiple access requests to obtain access addresses corresponding to the multiple access requests; determining whether the multiple access requests include multiple conflicting access requests whose access addresses point to the same target storage bank among the multiple storage banks; In response to the multiple access requests including the multiple conflicting access requests, determining a target conflicting access request from the multiple conflicting access requests, and performing a prefetch operation on the target conflicting access request, The prefetch operation includes prefetching target data in a target address corresponding to the access address of the target conflicting access request into a target register corresponding to the target storage body in the register module.

2. The data storage device according to claim 1, wherein The register module is configured as follows: In response to executing the prefetch operation, the target data in the target address is saved and the target address is recorded through the target register, so that the target thread corresponding to the target conflict access request reads the target data from the target register according to the target address.

3. The data storage device according to claim 2, wherein: The register module is further configured as: In response to the target thread finishing reading the target data in the target register, the target data and the recorded target address stored in the target register are cleared.

4. The data storage device according to claim 1, wherein The multiple conflicting access requests include a first conflicting access request corresponding to a first thread and a second conflicting access request corresponding to a second thread, the first conflicting access request is determined to be the target conflicting access request, and the first thread reads the target data from the target register and the second thread reads the first data from the target address of the target storage body in parallel.

5. The data storage device according to any one of claims 1 to 4, further comprising a scheduling module and a marking module, wherein: The scheduling module is configured to perform a scheduling operation on the multiple access requests; The marking module is configured to perform a marking operation on the target conflicting access request, so that the target conflicting access request skips the scheduling operation of the scheduling module.

6. The data storage device according to claim 5, further comprising: The arbitration module is configured to arbitrate the multiple access requests so that each memory bank responds to at most one access request at a time.

7. The data storage device according to claim 6, wherein: The arbitration module includes a first arbitration unit and a second arbitration unit. The first arbitration unit is configured to map the plurality of access requests to the plurality of memory banks according to the scheduling result of the scheduling operation, The second arbitration unit is configured to retrieve the first data corresponding to the multiple access requests from the multiple storage banks.

8. The data storage device according to claim 7, wherein: In response to executing the prefetch operation, the first arbitration unit is also configured to determine the target storage body in the storage body module to which the target conflicting access request is to be mapped, and the target storage body is configured to send the target data in the target address to the second arbitration unit, and the second arbitration unit is further configured to send the target data to the target register in the register module.

9. The data storage device according to claim 7, further comprising: a sorting module, connected to the second arbitration unit and the register module, configured to obtain the first data corresponding to the multiple access requests from the second arbitration unit and / or obtain the target data corresponding to the target conflicting access request from the register module, sort the format of the first data and / or the target data, and send the sorted first data and / or the target data to the general register connected to the data storage device.

10. The data storage device according to any one of claims 1 to 4, wherein: The N registers in the register module are configured to correspond to the same memory bank, where N is an integer greater than or equal to 1.

11. A data access method, comprising: Acquire multiple access requests from multiple threads for accessing a storage module, wherein the storage module includes multiple storages; Parsing the multiple access requests to obtain access addresses corresponding to the multiple access requests; determining whether the multiple access requests include multiple conflicting access requests whose access addresses point to the same target storage bank among the multiple storage banks; In response to the multiple access requests including the multiple conflicting access requests, determining a target conflicting access request from the multiple conflicting access requests, and performing a prefetch operation on the target conflicting access request, The prefetch operation includes prefetching target data in a target address corresponding to the access address of the target conflicting access request into a target register corresponding to the target storage body in a register module, and the register module includes multiple registers corresponding to the multiple storage bodies.

12. The data access method according to claim 11, further comprising: In response to executing the prefetch operation, the target data in the target address is saved and the target address is recorded through the target register, so that the target thread corresponding to the target conflict access request reads the target data from the target register according to the target address.

13. The data access method according to claim 12, further comprising: In response to the target thread finishing reading the target data in the target register, the target data and the recorded target address stored in the target register are cleared.

14. The data access method according to claim 11, wherein: The multiple conflicting access requests include a first conflicting access request corresponding to a first thread and a second conflicting access request corresponding to a second thread, the first conflicting access request is determined to be the target conflicting access request, and the first thread reads the target data from the target register and the second thread reads the first data from the target address of the target storage body in parallel.

15. The data access method according to any one of claims 11 to 14, further comprising: performing a scheduling operation on the multiple access requests; A marking operation is performed on the target conflicting access request so that the target conflicting access request skips the scheduling operation.

16. The data access method according to claim 15, further comprising: The multiple access requests are arbitrated so that each memory bank responds to at most one access request at a time.

17. The data access method according to claim 16, further comprising: mapping the plurality of access requests to the plurality of memory banks according to the scheduling result of the scheduling operation; The first data corresponding to the multiple access requests are retrieved from the multiple storage bodies.

18. An electronic device comprising the data storage device according to any one of claims 1 to 10.

19. The electronic device according to claim 18, further comprising a graphics processor, wherein The data storage device is provided on the stream multiprocessor of the graphics processor as a shared memory, The stream multiprocessor is further provided with a stream processor connected to the data storage device, and the stream processor is configured to send the multiple access requests to the data storage device.