Bidirectional storage distributed sorting method and system based on FPGA

By adopting a bidirectional storage distributed sorting method in FPGA and utilizing the bidirectional storage and iterative sorting of RAMX and RAMY, the problem of balancing time complexity and resource consumption is solved, and efficient data sorting is achieved.

CN120596058APending Publication Date: 2025-09-05SCHOOL OF INFORMATION & COMM TECH NAT UNIV OF DEFENSE TECH OF THE CHINESE PEOPLES LIBERATION ARMY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510640136.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing FPGA-based sorting schemes have difficulty balancing time complexity and resource consumption. Merge sorting is slow while distributed sorting consumes too much resources.

Method used

A bidirectional storage distributed sorting method is adopted. Two RAM blocks (RAMX and RAMY) are used in FPGA for bidirectional data storage and iterative sorting. The storage is judged according to the data bit width, and the storage intervals are switched by bit judgment during the iteration process until the target number of iterations is reached.

Benefits of technology

It improves the sorting speed and reduces resource consumption, especially when the data width is large. It has the advantages of both resource saving of merge sorting and speed of distributed sorting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120596058A_ABST
    Figure CN120596058A_ABST
Patent Text Reader

Abstract

The invention discloses a bidirectional storage distributed sorting method and system based on an FPGA, and the method comprises the steps: to-be-sorted data is written into a first storage block in a first order through a write address generation module, and the first storage block is RAMX; reading operation is executed from the first storage block RAMX, the first data is written into a second storage block according to a second sequence through bit judgment, and the second storage block is RAMY; reading operation is executed from the second storage block RAMY, and second data are written into the first storage block RAMX according to a third sequence through bit judgment to form bidirectional storage; and repeatedly executing a reading, storing and sorting process until a target number of iterations is reached, and reading a sorted result from the first storage block RAMX or the second storage block RAMY.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer data processing, and more specifically, to a bidirectional storage distributed sorting method and system based on FPGA. Background Art

[0002] Sorting is a basic data processing operation in the field of computer technology, and has very wide applications in data processing, databases, data compression, distributed computing and other fields.

[0003] Currently, most FPGA-based sorting solutions focus on merging algorithms, as they align closely with the "trading space for time" design philosophy, which is highly compatible with the FPGA's parallel processing model. However, merging algorithms are essentially comparison-based algorithms, sorting by repeatedly comparing key values. The lower bound of the time complexity of such algorithms is [unclear - likely a typo]. For example, the sorting method proposed in Chinese patent CN114416020A essentially expands the CPU's serial comparison method to an M-way parallel comparison, increasing the comparison speed by a factor of M. The FPGA-based bitonic sorting algorithm proposed in Chinese patent CN118192928A divides the data to be sorted into multiple subsequences of length 16, performs a first-level bitonic sort on each of these subsequences, and then sorts these subsequences in the order of monotonically increasing, monotonically decreasing, monotonically increasing, and monotonically decreasing before performing the next level of bitonic sorting. This step is repeated repeatedly until the entire sequence is sorted. This method is only applicable to sorting sequences of length [unclear - likely a typo] and has poor scalability. At the same time, the submodules need to be sorted according to certain rules between the two-level double-adjustment sorting modules, and the control logic is highly complex.

[0004] Compared to merge-based sorting, distributed sorting methods offer certain advantages in specific application scenarios. The basic principle of distributed sorting, such as radix / bucket sort, is to store data in different storage areas based on size and then sort the data based on the storage area. The sorting time is related to the data bit width, and for a k-bit data set, the time complexity is [ ]. Therefore, theoretically, radix / bucket sort is more advantageous when the sorting data meets the requirements. However, due to the unpredictable data distribution and range, the division of storage areas and the number of data in each storage area in radix / bucket sort cannot be determined. This places a high demand on on-chip storage and relies on large-scale storage to alleviate read and write storage overhead. Therefore, there is currently no research on distributed sorting acceleration for FPGAs. Summary of the Invention

[0005] In response to at least one defect or improvement need in the prior art, the present invention provides an FPGA-based bidirectional storage distributed sorting method and system, which solves the problem of FPGA-based sorting schemes that are difficult to balance time complexity and resource consumption. Data is bidirectionally stored according to data size, which improves sorting speed and reduces resource consumption.

[0006] To achieve the above-mentioned objectives, according to a first aspect of the present invention, a bidirectional storage distributed sorting method based on FPGA is provided, the method comprising: writing the data to be sorted into a first storage block in a first order through a write address generation module, wherein the first storage block is RAMX; performing a read operation from the first storage block RAMX, and writing the first data into a second storage block in a second order through bit judgment, wherein the second storage block is RAMY; performing a read operation from the second storage block RAMY, and writing the second data into the first storage block RAMX in a third order through bit judgment to form bidirectional storage; repeating the read storage sorting process until the target number of iterations is reached, and reading the sorted result from the first storage block or the second storage block.

[0007] In an exemplary embodiment, writing the to-be-sorted data into the first storage block in the first order by the write address generation module includes:

[0008] The RAM depth is the same as the length of the data to be sorted, and the RAM data bit width is n; the data to be sorted is stored in the first storage block RAMX, and the nth bit of the data to be sorted is judged by bit judgment when storing the data; the data with the nth bit being 0 and the data with the nth bit being 1 in the data to be sorted are respectively stored in two mutually isolated storage areas, wherein the storage areas are respectively the first storage area and the second storage area.

[0009] In an exemplary embodiment, after performing a read operation from the RAMX and writing the first data into the second storage block in the second order through bit judgment, the method further includes: traversing the data in the first storage block RAMX and judging the n-1th bit of the data; storing the data with the n-1th bit being 0 and the data with the n-1th bit being 1 in the first storage area in two mutually isolated third storage areas and fourth storage areas of the second storage block RAMY, respectively; and storing the data with the n-1th bit being 0 and the data with the n-1th bit being 1 in the second storage area in two mutually isolated fifth storage areas and sixth storage areas of the second storage block RAMY, respectively.

[0010] In an exemplary embodiment, after storing the data with the n-1th bit being 0 and the data with the n-1th bit being 1 in the second storage area in two mutually isolated fifth storage areas and sixth storage areas of the second storage block RAMY, respectively, the method further includes: traversing the data in the second storage block RAMY and judging the n-2th bit of the data; dividing the data in each storage area according to the value of the n-2th bit of the data, and storing them in the first storage block RAMX.

[0011] According to a second aspect of the present invention, there is also provided an FPGA-based bidirectional storage distributed sorting system, which applies the above-mentioned FPGA-based bidirectional storage distributed sorting method, including: a write address generation module for writing the data to be sorted into the RAM in a specific order; a RAM module including a first storage block RAMX and a second storage block RAMY isolated from each other, each storage block containing multiple isolated storage areas for storing data; a sorting module for reading from the first storage block RAMX, writing into the second storage block RAMY in a specific order through bit judgment, and then reading from the second storage block RAMY, writing into the first storage block RAMX in a specific order through bit judgment; an output module for outputting the sorted result from the first storage block RAMX or the second storage block RAMY after reaching the number of iterations.

[0012] According to a third aspect of the present invention, a bidirectional storage distributed sorting device based on FPGA is also provided, which includes: a RAM depth that is the same as the length of the data to be sorted, and a RAM data bit width of n; a first storage unit, used to store the data to be sorted in the first storage block RAMX, and to judge the nth bit of the data to be sorted by bit judgment when storing the data; a second storage unit, used to store data with the nth bit being 0 and data with the nth bit being 1 in the data to be sorted in two isolated storage areas, respectively, wherein the storage areas are the first storage area and the second storage area.

[0013] According to a fourth aspect of the present invention, a computer-readable storage medium is further provided, in which a computer program is stored, wherein the computer program is configured to execute the above-mentioned FPGA-based bidirectional storage distributed sorting method when running.

[0014] According to a fifth aspect of the present invention, an electronic device is also provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the FPGA-based bidirectional storage distributed sorting method through the computer program.

[0015] In general, the above technical solutions conceived by the present invention can achieve the following beneficial effects compared with the prior art:

[0016] The present invention provides an FPGA-based bidirectional storage distributed sorting method that performs bidirectional storage of data according to data size, thereby solving the problem in the prior art of the difficulty in balancing time complexity and resource consumption. For a data sequence with a bit width of N and a length of L, if distributed sorting is used, a storage area with a size of 2^N is required, while the present invention only requires a storage area with a depth of L. The advantage is very obvious when the data bit width is large, which is close to the resource consumption of merge-type sorting, but the sorting speed of the present invention can reach the level of distributed sorting. In summary, the present invention has the advantages of low resource consumption of the merge-type sorting method and high sorting speed of the distributed sorting method. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0018] Figure 1 A schematic flow chart of an optional FPGA-based bidirectional storage distributed sorting method provided in an embodiment of the present application;

[0019] Figure 2 An optional system structure diagram provided for an embodiment of the present application;

[0020] Figure 3 An optional test data schematic diagram provided for an embodiment of the present application;

[0021] Figure 4 A schematic diagram of an optional first iterative traversal process provided in an embodiment of the present application;

[0022] Figure 5 A schematic diagram of an optional second iterative traversal process provided in an embodiment of the present application;

[0023] Figure 6 A schematic diagram of an optional last three iterative traversal process provided in an embodiment of the present application;

[0024] Figure 7 This is a schematic diagram of the traversal process when the present invention uses double resources for iterative traversal;

[0025] Figure 8 A schematic diagram of the structure of an optional FPGA-based bidirectional storage distributed sorting device provided in an embodiment of the present application;

[0026] Figure 9 A schematic structural diagram of an optional electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0027] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below may be combined with each other as long as they do not conflict with each other.

[0028] The terms "first," "second," "third," and the like in the specification and claims of this application and the accompanying drawings are used to distinguish between different objects, not to describe a particular order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements, but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or apparatus.

[0029] According to one aspect of the embodiment of the present application, a bidirectional storage distributed sorting method based on FPGA is provided. Figure 1 The present invention describes an FPGA-based bidirectional storage distributed sorting method provided in an embodiment of the present application.

[0030] Figure 1 This is a flow chart of an optional FPGA-based bidirectional storage distributed sorting method provided in an embodiment of the present application, such as Figure 1 As shown, the process of the method may include the following steps:

[0031] S102, the data to be sorted is written into a first storage block in a first order by a write address generation module, wherein the first storage block is RAMX;

[0032] S104, performing a read operation from the first storage block RAMX, and writing the first data into the second storage block in a second order through bit determination, wherein the second storage block is RAMY;

[0033] S106, performing a read operation from the second storage block RAMY, and writing the second data into the first storage block RAMX in a third order after bit determination to form bidirectional storage;

[0034] S108 , repeatedly executing the read storage sorting process until the target number of iterations is reached, and then reading the sorted result from the first storage block RAMX or the second storage block RAMY.

[0035] The bidirectional storage distributed sorting method based on Field Programmable Gate Array (FPGA) in this application can be applied to the scenario where data is stored in different storage areas according to the data size and then the data is sorted according to the storage area. Figure 2 and Figure 3 As shown, the FPGA-based bidirectional storage distributed sorting method may include the following steps: first, two storage blocks may be opened up, including a first storage block and a second storage block (for example, RAMX and RAMY). Here, the RAM depth is L, which is consistent with the length of the data to be sorted, and the RAM data bit width is n.

[0036] It's important to note that in the field of FPGAs and embedded systems, RAMX typically refers to extended RAM (random access memory). In some FPGA designs, RAMX may be used to represent an extended storage unit, distinct from the basic RAM (internal RAM), used to store additional data or program code. For example, in some FPGA implementations of 8051 microcontrollers, RAMX may be used to expand the capacity of the internal RAM. The secondary memory block, RAMY, is a secondary memory in addition to the primary memory and is sometimes referred to as the secondary memory block.

[0037] For example, two storage blocks RAMX and RAMY can be opened. The data to be sorted first passes through the write address generation module and is written into RAMX in a specific order. Then it is read from RAMX and written into RAMY in a specific order through bit judgment. Then it is read from RAMY and written into RAMX in a specific order through bit judgment. This process is repeated until the number of iterations is reached, and the sorted result is read from RAMX or RAMY.

[0038] Specifically, the data to be sorted is stored in RAMX, and the nth bit of the data to be sorted is judged when storing the data, and the data with the nth bit being "0" and the data with the nth bit being "1" are stored in two isolated storage areas A and B respectively. Traverse the data in RAMX, judge the n-1th bit of the data, and store the data with the n-1th bit being "0" and the data with the n-1th bit being "1" in storage area A in two mutually isolated storage areas A1 and storage area A2 of RAMY respectively; store the data with the n-1th bit being "0" and the data with the n-1th bit being "1" in storage area B in two mutually isolated storage areas B1 and storage area B2 of RAMY respectively; traverse the data in RAMY, judge the n-2th bit of the data, and similarly to step S3, further divide the data in each storage area in step S3 according to the value of the n-2th bit of the data and store them in RAMX; repeat the above steps, iteratively updating the data in RAMX and RAMY according to the ni (i=1, ..., n-1)th bit of the data in RAMX and RAMY in turn. When i=n-1, the sequence obtained by the last update is the sorted result.

[0039] Through the above steps S102 to S108, the data to be sorted is written into the first storage block in the first order through the write address generation module, wherein the first storage block is RAMX; a read operation is performed from the first storage block RAMX, and the first data is written into the second storage block in the second order through bit judgment, wherein the second storage block is RAMY; a read operation is performed from the second storage block RAMY, and the second data is written into the first storage block RAMX in the third order through bit judgment to form bidirectional storage; the read storage sorting process is repeated until the target number of iterations is reached, and the sorted result is read out from the first storage block RAMX or the second storage block RAMY, thereby solving the problem of balancing time complexity and resource consumption in the FPGA-based sorting scheme, and bidirectionally storing data according to data size, thereby improving the sorting speed and reducing resource consumption.

[0040] In an exemplary embodiment, writing the to-be-sorted data into the first storage block in the first order by the write address generation module includes:

[0041] The RAM depth is the same as the length of the data to be sorted, and the RAM data bit width is n;

[0042] S11, storing the data to be sorted into the first storage block RAMX, and performing bit judgment on the nth bit of the data to be sorted during data storage;

[0043] S12, storing the data with the nth bit being 0 and the data with the nth bit being 1 in the data to be sorted in two mutually isolated storage areas, respectively, wherein the storage areas are a first storage area and a second storage area.

[0044] In the embodiment of the present application, a set of unsigned data with a bit width of 5 and a length of 19 is taken as an example to explain the implementation of the present invention in detail. Figure 3 As shown, two RAMX and RAMY with a width of 5 and a depth of 19 are generated.

[0045] The sorting method process is as follows Figure 2 As shown, the data to be sorted first passes through the write address generation module and is written to RAMX in a specific order. It is then read from RAMX and written to RAMY in a specific order through bit judgment. The data is then read from RAMY and written to RAMX in a specific order through bit judgment. This process is repeated until the number of iterations is reached, and the sorted result is read from RAMX or RAMY.

[0046] See also Figure 4 The first step is to store the test data into RAMX according to the 5th bit of the test data. The specific storage method is that if the 5th bit of the data is 0, the data is stored in the corresponding address of RAMX in the order of 0, 1, 2, ...; if the 5th bit of the data is 1, the data is stored in the corresponding address of RAMX in the order of 18, 17, 16, ...

[0047] In an exemplary embodiment, after performing a read operation from the RAMX and writing the first data into the second storage block in the second order through bit determination, the method further includes:

[0048] S21, traverse the data in the first storage block RAMX and make a decision on the n-1th bit of the data;

[0049] S22, storing the data with the n-1th bit being 0 and the data with the n-1th bit being 1 in the first storage area in two mutually isolated third and fourth storage areas of the second storage block RAMY respectively;

[0050] S23 , storing the data with the n-1th bit being 0 and the data with the n-1th bit being 1 in the second storage area in two mutually isolated fifth and sixth storage areas of the second storage block RAMY respectively.

[0051] In the examples of this application, see Figure 5, after the first traversal, the data is divided into two parts. The MIN block is stored at addresses 0 - 12, and the MAX block is stored at addresses 13 - 18. All the data in the MAX block is greater than the data in the MIN block. Next, a second traversal is performed to judge the 4th bit of the data. For the data in the MIN block, if the 4th bit is 0, the data is stored in the corresponding addresses in the order of 0, 1, 2, …; if the 4th bit of the data is 1, the data is stored in the corresponding addresses in the order of 12, 11, 10, …. For the data in the MAX block, if the 4th bit is 0, the data is stored in the corresponding addresses in the order of 13, 14, 15, … in RAMY; if the 4th bit of the data is 1, the data is stored in the corresponding addresses in the order of 18, 17, 16, … in RAMY.

[0052] In an exemplary embodiment, after storing the data with the (n - 1)th bit being 0 and the data with the (n - 1)th bit being 1 in the second storage area in two mutually isolated fifth and sixth storage areas of the second storage block RAMY respectively, the above method further includes:

[0053] S31, traverse the data in the second storage block RAMY and judge the (n - 2)th bit of the data;

[0054] S32, after dividing the data in each storage area according to the value of the (n - 2)th bit of the data, store it in the first storage block RAMX.

[0055] In the embodiment of the present application, refer to Figure 6 , after the second traversal, the data is divided into four parts and stored in RAMX as area A with addresses 0 - 5, area B with addresses 6 - 12, area C with addresses 13 - 15, and area D with addresses 16 - 18. Among them, the data in area A < the data in area B < the data in area C < the data in area D. Subsequently, through similar three traversal processes, the sorted data can be obtained.

[0056] As can be seen from the above embodiments, for a sequence with a bit width of k and a data length of N, the time complexity of the present invention is O(N×k). When the sorted data satisfies k < log2N, it is better than the merge sort algorithm and has the same time complexity as the distributed sort, but only two RAMs with a bit width of k and a depth of N are consumed, greatly reducing the resources consumed compared to the distributed sorting method.

[0057] Furthermore, if true dual-port RAM is used, data can be read and written simultaneously from both ports, doubling the iteration speed and reducing the time complexity to O(Nk / 2). If double the resource consumption, such as using two 2L-deep RAMs, two bits of data can be determined at a time. For example, two 2L-deep memory areas, RAMX and RAMY, are generated. Data A, whose top two bits are "00," is stored in RAMX in the order 0, 1, 2, ...; data B, whose top two bits are "01," is stored in the order N-1, N-2, N-3, ...; data C, whose top two bits are "10," is stored in the order N, N+1N+2, ...; and data D, whose top two bits are "11," is stored in the order 2N-1, 2N-2, 2N-3, .... Assume that the highest two bits are "00", "01", "10", and "11", and the number of data is m1, m2, m3, and m4 (m1+m2+m3+m4=N).

[0058] During the first traversal, for data A, if its (N-2, N-3)th bit data is "00", it is stored in the order of 0, 1, 2, ...; if it is "01", it is stored in the order of m1-1, m1-1, m1-2, ...; if it is "01", it is stored in the order of m1, m1+1, m1+2, ...; if it is "11", it is stored in the order of 2m1-1, 2m1-1, 2m1-2, ...

[0059] For data B, the iterative traversal process is as follows: Figure 7 As shown, if the (N-2, N-3)th bit data is "00", it is stored in the order of 2m1, 2m1+1, 2m1+2...; if it is "01", it is stored in the order of 2m1+m2-1, 2m1+m2-1, 2m1+m2-2, ...; if it is "01", it is stored in the order of 2m1+m2, 2m1+m2+1, 2m1+m2+2, ...; if it is "11", it is stored in the order of 2m1+2m2-1, 2m1+2m2-1, 2m1+2m2-2, .... The storage addresses of C data and D data can also be deduced according to the rule.

[0060] As can be seen above, if double the resources are used, one iteration can determine two bits of data, halving the number of iterations required for sorting. Furthermore, if this doubling of resource consumption is achieved by generating two equal-sized RAMs, the speed of each iteration can also be doubled. Theoretically, using four RAMs for sorting can result in a time complexity of . Similarly, further increasing the storage resources can further shorten the sorting time.

[0061] According to another aspect of an embodiment of the present application, a bidirectional storage distributed sorting system is provided that applies the above-mentioned bidirectional storage distributed sorting method based on FPGA. The bidirectional storage distributed sorting method based on FPGA is applied, and includes:

[0062] A write address generation module is used to write the data to be sorted into the RAM in a specific order;

[0063] The RAM module includes a first storage block RAMX and a second storage block RAMY isolated from each other, each storage block including a plurality of isolated storage areas for storing data;

[0064] a sorting module, configured to read from the first storage block RAMX, write into the second storage block RAMY in a specific order through bit judgment, and then read from the second storage block RAMY, write into the first storage block RAMX in a specific order through bit judgment;

[0065] The output module is used to output the sorted result from the first storage block RAMX or the second storage block RAMY after the number of iterations is reached.

[0066] According to another aspect of the embodiments of the present application, a bidirectional storage distributed sorting device for implementing the above-mentioned FPGA-based bidirectional storage distributed sorting method is also provided. Figure 8 is a structural diagram of an optional FPGA-based bidirectional storage distributed sorting device according to an embodiment of the present application, such as Figure 8 As shown, the device may include:

[0067] The RAM depth is the same as the length of the data to be sorted, and the RAM data bit width is n;

[0068] A first storage unit 802 is configured to store the data to be sorted into the first storage block RAMX, and to perform bit determination on the nth bit of the data to be sorted when storing the data;

[0069] The second storage unit 804 is used to store the data with the nth bit being 0 and the data with the nth bit being 1 in the data to be sorted in two isolated storage areas, respectively, wherein the storage areas are the first storage area and the second storage area.

[0070] Through the above module, the data to be sorted is written into the first storage block in the first order through the write address generation module, wherein the first storage block is RAMX; a read operation is performed from the first storage block RAMX, and the first data is written into the second storage block in the second order through bit judgment, wherein the second storage block is RAMY; a read operation is performed from the second storage block RAMY, and the second data is written into the first storage block RAMX in the third order through bit judgment to form bidirectional storage; the read storage sorting process is repeated until the target number of iterations is reached, and the sorted result is read out from the first storage block RAMX or the second storage block RAMY, thereby solving the problem of balancing time complexity and resource consumption in the FPGA-based sorting scheme, and bidirectionally storing data according to data size, thereby improving the sorting speed and reducing resource consumption.

[0071] In an exemplary embodiment, the apparatus further comprises:

[0072] A first decision unit, configured to traverse the data in the first storage block RAMX and make a decision on the n-1th bit of the data;

[0073] a third storage unit, configured to store the data with the n-1th bit being 0 and the data with the n-1th bit being 1 in the first storage area in two mutually isolated third and fourth storage areas of the second storage block RAMY, respectively;

[0074] The fourth storage unit is used to store the data with the n-1th bit being 0 and the data with the n-1th bit being 1 in the second storage area in two mutually isolated fifth and sixth storage areas of the second storage block RAMY respectively.

[0075] In an exemplary embodiment, the apparatus further comprises:

[0076] A second decision unit, configured to traverse the data in the second storage block RAMY and make a decision on the n-2th bit of the data;

[0077] The fifth storage unit is used to divide the data in each storage area according to the value of the n-2th bit of the data and store the data in the first storage block RAMX.

[0078] It should be noted here that the examples and scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the contents disclosed in the above embodiments. It should be noted that the above modules as part of the device can run in a hardware environment, can be implemented by software, and can also be implemented by hardware, where the hardware environment includes a network environment.

[0079] According to another aspect of the embodiments of the present application, a storage medium is further provided. Optionally, in this embodiment, the storage medium can be used to execute the program code of any of the above-mentioned bidirectional storage distributed sorting methods based on FPGA in the embodiments of the present application.

[0080] Optionally, in this embodiment, the storage medium is configured to store program codes for executing the following steps:

[0081] S1, the data to be sorted is written into the first storage block in a first order through the write address generation module, wherein the first storage block is RAMX;

[0082] S2, performing a read operation from the first storage block RAMX, and writing the first data into the second storage block in a second order through bit judgment, wherein the second storage block is RAMY;

[0083] S3, performing a read operation from the second storage block RAMY, and writing the second data into the first storage block RAMX in a third order after bit judgment to form bidirectional storage;

[0084] S4, repeatedly executing the read storage sorting process until the target number of iterations is reached, and then reading the sorted result from the first storage block RAMX or the second storage block RAMY.

[0085] Optionally, the specific examples in this embodiment may refer to the examples described in the above embodiments, which will not be described in detail in this embodiment.

[0086] Among them, computer-readable storage media may include, but are not limited to, any type of disk, including floppy disks, optical disks, DVDs, CD-ROMs, microdrives and magneto-optical disks, ROMs, RAMs, EPROMs, EEPROMs, DRAMs, VRAMs, flash memory devices, magnetic or optical cards, nanosystems (including molecular memory ICs), or any type of medium or device suitable for storing instructions and / or data.

[0087] According to another aspect of the embodiments of the present application, an electronic device for implementing the above-mentioned FPGA-based bidirectional storage distributed sorting method is also provided. The electronic device can be a server, a terminal, or a combination thereof.

[0088] Figure 9 is a schematic structural diagram of an optional electronic device according to an embodiment of the present application, such as Figure 9 As shown, it includes a processor 902, a communication interface 904, a memory 906 and a communication bus 908, wherein the processor 902, the communication interface 904, and the memory 906 communicate with each other via the communication bus 908, wherein,

[0089] Memory 906, for storing computer programs;

[0090] The processor 902 is configured to execute the computer program stored in the memory 906 to implement the following steps:

[0091] S1, the data to be sorted is written into the first storage block in a first order through the write address generation module, wherein the first storage block is RAMX;

[0092] S2, performing a read operation from the first storage block RAMX, and writing the first data into the second storage block in a second order through bit judgment, wherein the second storage block is RAMY;

[0093] S3, performing a read operation from the second storage block RAMY, and writing the second data into the first storage block RAMX in a third order after bit judgment to form bidirectional storage;

[0094] S4, repeatedly executing the read storage sorting process until the target number of iterations is reached, and then reading the sorted result from the first storage block RAMX or the second storage block RAMY.

[0095] Optionally, the communication bus may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. The communication bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 9 The communication interface is used for communication between the electronic device and other devices.

[0096] The memory may include RAM, or may include non-volatile memory, such as at least one disk memory. Alternatively, the memory may also be at least one storage device located away from the aforementioned processor.

[0097] As an example, the memory 906 may include, but is not limited to, the first storage unit 802 and the second storage unit 804 in the FPGA-based bidirectional storage distributed sorting device. Furthermore, the memory 906 may also include, but is not limited to, other module units in the FPGA-based bidirectional storage distributed sorting device, which will not be described in detail in this example.

[0098] The above-mentioned processor can be a general-purpose processor, including but not limited to: CPU (Central Processing Unit), NP (Network Processor), etc.; it can also be DSP (Digital Signal Processing), ASIC (Application Specific Integrated Circuit), FPGA (Field-Programmable Gate Array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0099] Optionally, the specific examples in this embodiment may refer to the examples described in the above embodiments, and this embodiment will not be described in detail here.

[0100] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.

[0101] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0102] In the several embodiments provided in this application, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some service interface, and the indirect coupling or communication connection of the device or unit can be electrical or other forms.

[0103] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0104] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0105] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a memory, including a number of instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned memory includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0106] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable memory, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0107] The above is only an exemplary embodiment of the present disclosure and cannot be used to limit the scope of the present disclosure. That is, any equivalent changes and modifications made according to the teachings of the present disclosure are still within the scope of the present disclosure. After considering the specification and practicing the disclosure herein, those skilled in the art will easily think of the implementation scheme of the present disclosure. This application is intended to cover any variation, use or adaptation of the present disclosure, which follows the general principles of the present disclosure and includes common knowledge or customary technical means in the art that are not recorded in the present disclosure. The description and examples are to be regarded as exemplary only, and the scope and spirit of the present disclosure are defined by the claims.

[0108] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0109] It will be easily understood by those skilled in the art that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A bidirectional storage distributed sorting method based on FPGA, characterized in that: include: The data to be sorted is written into a first storage block in a first order by a write address generation module, wherein the first storage block is RAMX; Perform a read operation from the first storage block RAMX, and write the first data into the second storage block in a second order through bit judgment, wherein the second storage block is RAMY; Perform a read operation from the second storage block RAMY, and write the second data into the first storage block RAMX in a third order after bit judgment to form a bidirectional storage; After repeatedly executing the read storage sorting process until the target number of iterations is reached, the sorted result is read out from the first storage block RAMX or the second storage block RAMY.

2. The FPGA-based bidirectional storage distributed sorting method according to claim 1, wherein: The data to be sorted is written into the first storage block in a first order by the write address generation module, comprising: The RAM depth is the same as the length of the data to be sorted, and the RAM data bit width is n; Storing the data to be sorted in the first storage block RAMX, and judging the nth bit of the data to be sorted by bit judgment when storing the data; The data with the nth bit being 0 and the data with the nth bit being 1 in the data to be sorted are respectively stored in two mutually isolated storage areas, wherein the storage areas are respectively a first storage area and a second storage area.

3. The FPGA-based bidirectional storage distributed sorting method according to claim 1, wherein: After performing a read operation from the RAMX and writing the first data into the second storage block in a second order through bit judgment, the method further includes: Traversing the data in the first storage block RAMX and making a judgment on the n-1th bit of the data; storing the data with the n-1th bit being 0 and the data with the n-1th bit being 1 in the first storage area in two mutually isolated third storage areas and fourth storage areas of the second storage block RAMY respectively; The data with the n-1th bit being 0 and the data with the n-1th bit being 1 in the second storage area are respectively stored in the second storage block RAMY and the second storage area being isolated from each other.

4. The FPGA-based bidirectional storage distributed sorting method according to claim 3, wherein: After storing the data with the n-1th bit being 0 and the data with the n-1th bit being 1 in the second storage area in the two mutually isolated fifth storage areas and sixth storage areas of the second storage block RAMY respectively, the method further includes: Traversing the data in the second storage block RAMY, and making a judgment on the n-2th bit of the data; The data in each storage area is divided according to the value of the n-2th bit of the data and then stored in the first storage block RAMX.

5. A bidirectional storage distributed sorting system based on FPGA, applying the bidirectional storage distributed sorting method based on FPGA according to any one of claims 1 to 4, comprising: A write address generation module is used to write the data to be sorted into the RAM in a specific order; The RAM module includes a first storage block RAMX and a second storage block RAMY isolated from each other, each storage block including a plurality of isolated storage areas for storing data; a sorting module, configured to read from the first storage block RAMX, write into the second storage block RAMY in a specific order through bit judgment, and then read from the second storage block RAMY, write into the first storage block RAMX in a specific order through bit judgment; The output module is used to output the sorted result from the first storage block RAMX or the second storage block RAMY after the number of iterations is reached.

6. A bidirectional storage distributed sorting device based on FPGA, characterized in that: include: The RAM depth is the same as the length of the data to be sorted, and the RAM data bit width is n; a first storage unit, configured to store the data to be sorted into the first storage block RAMX, and to judge the nth bit of the data to be sorted by bit judgment when storing the data; The second storage unit is used to store the data with the nth bit being 0 and the data with the nth bit being 1 in the data to be sorted in two mutually isolated storage areas, wherein the storage areas are the first storage area and the second storage area respectively.

7. The FPGA-based bidirectional storage distributed sorting device according to claim 6, characterized in that: The device further comprises: A first decision unit, configured to traverse the data in the first storage block RAMX and make a decision on the n-1th bit of the data; a third storage unit, configured to store the data with the n-1th bit being 0 and the data with the n-1th bit being 1 in the first storage area in two mutually isolated third and fourth storage areas of the second storage block RAMY, respectively; The fourth storage unit is used to store the data with the n-1th bit being 0 and the data with the n-1th bit being 1 in the second storage area in two mutually isolated fifth and sixth storage areas of the second storage block RAMY respectively.

8. The FPGA-based bidirectional storage distributed sorting device according to claim 7, characterized in that: The device further comprises: A second decision unit, configured to traverse the data in the second storage block RAMY and make a decision on the n-2th bit of the data; The fifth storage unit is used to divide the data in each storage area according to the value of the n-2th bit of the data and store the data in the first storage block RAMX.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein the program executes the method according to any one of claims 1 to 6 when executed.

10. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 6 through the computer program.

Citation Information

Patent Citations

  • Quick sorting method and device based on FPGA (Field Programmable Gate Array)

    CN114416020A

  • FPGA (Field Programmable Gate Array)-based dual-tone sorting method and FPGA

    CN118192928A