A GPU parallel intelligent search method, medium and device for wide-area spatial railway lines
By parallelizing the grid scanning of railway line search on the GPU, using the parallel processing capabilities of the GPU and the coordinated control of the CPU, the performance bottleneck problem of railway line search algorithms in wide-area space and fine grids in the prior art is solved, and more efficient line search and better solution discovery is achieved.
Patent Information
- Application Number
- CN202510179442.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-19
AI Technical Summary
When existing railway line search algorithms conduct path searches in wide-area space and fine grids, they face performance bottlenecks, resulting in too long search time, and simplified grid processing may ignore better line solutions, especially in complex terrain areas.
The GPU parallel intelligent search method is adopted to establish a comprehensive geographic information model, divide the grid hierarchy, and scan it in parallel in the GPU to update the distance map and find feasible line paths. This method utilizes the parallel processing capabilities of the GPU and combines the collaborative computing architecture controlled by CPU to achieve efficient application of distance conversion method.
It significantly improves the computing efficiency of railway line search, can process more refined geographic information model data, reduces search time, and increases the probability of discovering better line solutions, especially in complex terrain areas.
Smart Images

Figure CN119671005B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of railway line intelligent search, and in particular to a wide-area spatial railway line GPU parallel intelligent search method, medium and equipment. Background Art
[0002] Railway lines play a vital role in railway construction. As the leader and foundation of railway survey and design, line selection design is a core work in railway construction that involves a wide range of areas and is highly systematic. It has a decisive impact on the difficulty of the project, the size of the project investment, and the safety of construction and operation. Traditional manual line selection is limited by time, energy and equipment. Designers can only select a small number of schemes for detailed design based on experience, and it is difficult to ensure the optimal scheme. Line intelligent search and optimization technology can use computers to automatically complete the three-dimensional space line search and coordinated layout of structures for railways, and generate line plans that meet various constraints and have the optimal objective function. It effectively solves the problems of limited solutions, long decision-making cycle, single evaluation index, and high labor intensity in existing manual line selection design, and significantly improves design efficiency and quality.
[0003] At present, an effective route search method is the generalized distance transform method. Distance transform is a common transformation method in image science. It can convert a bitmap into a distance map. The distance map records the shortest distance from each pixel to the target pixel. According to the distance map, the shortest path to the target point can be generated. The meaning of distance is extended to consider the comprehensive engineering cost including roads, bridges and tunnels to form a generalized distance transform method. This method can automatically generate a large number of route paths connecting the start and end points in a complex mountainous environment. However, in practical applications, it is often necessary to search for paths in a wide spatial range and fine grids. In this case, the existing algorithms face serious performance bottlenecks. Even with high-performance computing equipment, a single path search may still take too long. In order to shorten the search time, the existing algorithms often use a simplified grid processing method, that is, to increase the grid size by pre-screening the grid data. Although this improves the calculation efficiency to a certain extent, it also has a greater probability of ignoring better route solutions, especially in complex terrain areas such as mountainous areas.
[0004] Therefore, there is an urgent need to propose a railway line search method that has high computational efficiency and can process more sophisticated geographic information model data to solve the problems existing in the prior art. Summary of the invention
[0005] The purpose of the present invention is to provide a wide-area spatial railway line GPU parallel intelligent search method, and its specific technical solution is as follows:
[0006] A wide-area spatial railway line GPU parallel intelligent search method comprises the following steps:
[0007] S1. Establish a comprehensive geographic information model and initialize data;
[0008] S2, divide the grid into levels;
[0009] S3, parallelize the grid scanning in the GPU to update the distance map and find a feasible line path;
[0010] S4. Determine whether the distance value in the distance map has changed. If so, return to S3; otherwise, output the final distance map data.
[0011] Preferably, the specific steps of S1 include:
[0012] S1-1, establishing a comprehensive geographic information model array of grid data, wherein the grid data includes elevation, grid type, fill and cut volume, bridge and tunnel, land acquisition unit price and control factor information;
[0013] S1-2, initialize the distance map array;
[0014] S1-3, initialize the scan control array;
[0015] S1-4, initialize the tag variable;
[0016] The comprehensive geographic information model array is stored in the GPU global memory.
[0017] Preferably, the distance map data in the distance map array includes a generalized distance value and a relative row offset and a relative column offset of a next grid on the shortest path from the current grid to the target point.
[0018] Preferably, the grid level division of S2 comprises the following steps:
[0019] Divide the grid data in S1 into blocks;
[0020] Each block is further divided into thread blocks;
[0021] Each thread block is responsible for processing A piece of grid data of size;
[0022] Based on the known size of the grid data, the block size, and the thread block size, the number of blocks contained in the grid data is calculated using the following formula:
[0023] ;
[0024] ;
[0025] Where: Indicates the total number of rows of blocks into which the grid data is divided. Indicates the total number of columns into which the grid data is divided; is the total number of rows of grid data, is the total number of columns of grid data; is the total number of rows or columns of grid data contained in the thread block; is the total number of rows or columns of thread blocks contained in the block.
[0026] Preferably, the S3 specifically includes the following steps:
[0027] S3-1, checking whether the current block needs to update the distance map data according to the scan control array;
[0028] S3-2, copy the local grid data and local distance map data corresponding to each block from the global memory to the shared memory where the block is located, and convert the global index into a local index during the copying;
[0029] S3-3, create a Boolean variable in the shared memory to control the scan loop of S3-4;
[0030] S3-4, all thread blocks continuously traverse and scan the grid within their range to find feasible line paths;
[0031] S3-5. Detect whether the distance map data changes during the scanning process. If changed, copy the local distance map data in the shared memory to the global memory.
[0032] Preferably, S3-1 specifically includes:
[0033] Create a size The two-dimensional Boolean array of is used as the scan control array and stored in the global memory of the GPU;
[0034] The scan control array is used to indicate whether the current block needs to be grid scanned to reduce the generalized distance value. The expression is as follows:
[0035] ;
[0036] Where: Indicates the scan control array;
[0037] The scan control array optimizes the calculation process by marking which blocks need further processing. When the distance map data in a block changes, the Ctrl value of the block and its adjacent blocks will be set to true, triggering the next round of grid scanning and updating the distance map.
[0038] Preferably, before the grid scanning in S3-4 starts, a local Boolean variable that can only be accessed within the thread block is defined, and if the thread block modifies the distance map data in the local shared memory during the scanning process, the variable is assigned a value of true, otherwise it defaults to false;
[0039] The steps for detecting whether the distance map data has changed during the scanning process in S3-5 are as follows: after the scanning is completed, if the local Boolean variable is true, the mark variable in the global memory is set to true using the CAS operation; if the mark variable is false, it means that all thread blocks in all blocks have not changed the distance map data.
[0040] Preferably, S4 includes:
[0041] S4-1, update the scan control array;
[0042] After each round of S3, the scan control array of each thread block will be updated; if the distance map data of the block in this round of search is changed, the scan control array value at the same position is true, then the scan control array values corresponding to the eight adjacent blocks around it will be set to true in the next round of search;
[0043] S4-2, check whether the distance value of this round of scanning has been modified. If it has been modified, return to S3; if it has not been modified, the algorithm ends;
[0044] The distance map data used by each block when performing grid scanning to search for paths comes from the distance map data of the previous round of scanning. By executing two rounds of S3 steps, the shortest distance from the adjacent block to the target point is propagated to the current block; after a maximum of In step S3 of the round, the distance map data of each block is no longer changed. At this time, what is stored in the distance map is the shortest distance from all grid points to the target point.
[0045] The application of the technical solution of the present invention has the following beneficial effects:
[0046] A GPU parallel intelligent search method for wide-area space railway lines includes the following steps: S1, gridding the line selection area, establishing a comprehensive geographic information model based on the grid and initializing data; S2, hierarchical division of the grid; S3, parallelizing grid scanning in the GPU to update the distance map and find a feasible line path; S4, judging whether the distance value in the distance map has changed, if changed, returning to S3, otherwise outputting the final distance map data. The present invention utilizes the powerful parallel processing capability of the GPU by hierarchically dividing a large number of grids, performing serial calculations in each thread block area, and parallel calculations of all thread blocks, and adopts a collaborative computing architecture of GPU computing and CPU control, so as to realize the efficient application of the distance transformation method in wide-area line selection space and fine grids.
[0047] The present invention also provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the aforementioned wide-area spatial railway line GPU parallel intelligent search method is implemented.
[0048] The present invention also provides an electronic device, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the aforementioned wide-area spatial railway line GPU parallel intelligent search method by executing the executable instructions.
[0049] In addition to the above-described purposes, features and advantages, the present invention has other purposes, features and advantages. The present invention will be further described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0051] Figure 1 is a flow chart of a wide-area spatial railway line GPU parallel intelligent search method of an embodiment;
[0052] Figure 2 This is a schematic diagram of grid level division;
[0053] Figure 3 To parallelize the grid scanning flow chart;
[0054] Figure 4 This is a schematic diagram of memory copy;
[0055] FIG5(a) is a schematic diagram of the first round of inter-block distance propagation, FIG5(b) is a schematic diagram of the second round of inter-block distance propagation, and FIG5(c) is a schematic diagram of the third round of inter-block distance propagation. DETAILED DESCRIPTION
[0056] The embodiments of the present invention are described in detail below with reference to the accompanying drawings. Example
[0057] The embodiment of the present invention is to Perform fast line path search for the line selection area, such as Figure 1 As shown, a wide-area spatial railway line GPU parallel intelligent search method includes the following steps:
[0058] S1. Establish a comprehensive geographic information model, which specifically includes the following steps:
[0059] S1-1. Establishing a comprehensive geographic information model array for grid data , which includes but is not limited to the following grid data: grid elevation, grid type, fill and cut, bridge and tunnel, land acquisition unit price, control factor information, etc.
[0060] S1-2. Initialize the distance map array The distance graph data includes the generalized distance value DT and the relative row offset of the next connected grid in the shortest path from the current grid to the target point. and column offset Etc. Before starting the calculation, all distance values in the distance map need to be set to a specific initial value. The expression is as follows:
[0061] ;
[0062] Where: Indicates the distance map is located at The generalized distance value of the grid, where the target point refers to the starting point or end point of the line.
[0063] S1-3. Initialize the scan control array , the initial values are all false.
[0064] S1-4. Initialize the flag variable flag, and the initial value is false.
[0065] In addition, before using the GPU for fast route search, given the relatively low data transfer rate between the CPU and the GPU, it is necessary to copy the comprehensive geographic information model array and the initialized distance map array prepared in step S1 to the GPU's global memory in one go. Once the data is copied, all subsequent related computing tasks can be completely executed by the GPU autonomously without the need for frequent data exchange with the CPU. This measure improves the time efficiency of the entire processing flow and reduces the delay caused by data communication.
[0066] S2, divide the grid into levels, see Figure 2 , including the following steps:
[0067] Divide the grid data mentioned in step S1 into A larger block (Block);
[0068] Each larger block is further divided into A smaller thread block (Thread);
[0069] Each thread block is responsible for processing A piece of grid data of size.
[0070] According to the above division levels, the grid data size is known, and the For 8, is 32, the block size of the grid data can be calculated by the following formula:
[0071] ;
[0072] ;
[0073] Where: and Respectively represent the total number of rows and the total number of columns of the blocks into which the grid data read from step S1 is divided; and are the total number of rows and columns of the grid data read in step S1 respectively; The total number of rows (columns) of grid data contained in the thread block (Thread); The total number of rows (columns) of thread blocks (Thread) contained in each block (Block).
[0074] Through this step, the grid is decomposed into several small areas. In the same round of grid scanning, all thread blocks in each block share the same local grid data, and at the same time, each thread block performs calculation operations simultaneously.
[0075] S3, parallelize the grid scanning in the GPU to update the distance map and find a feasible line path, see Figure 3 , specifically:
[0076] A step contains multiple sub-steps. In each block, all thread blocks execute all sub-steps in parallel. After each sub-step is executed, thread synchronization operations are performed to ensure that all thread blocks in the block have completed the current sub-step before entering the next sub-step.
[0077] S3-1, according to the scanning control array Check whether the distance map of the current block needs to be updated, specifically:
[0078] Step S1-3 establishes the scan control array , the purpose of this array is to indicate whether the current block needs to be grid scanned to reduce the generalized distance value. Specifically, its value is defined as follows:
[0079] ;
[0080] This scan control array optimizes the calculation process by marking which blocks need further processing. When the distance value in a block changes, the Ctrl value of this block and its neighboring blocks will be set to true, triggering the next round of grid scanning and updating the distance map. This mechanism ensures that only those blocks that may affect the final result are processed, thus avoiding unnecessary calculations.
[0081] When S3 is executed for the first time, the Ctrl array only contains the block where the target point is located. The corresponding value is true, and the others are false, so only The block will continue to execute the subsequent sub-steps. In the subsequent repeated execution of S3, only the block corresponding to the position where the Ctrl array value is true will execute the subsequent sub-steps of S3.
[0082] S3-2, the local grid data corresponding to each block is transferred from the global memory Local distance map data Copy to the high-speed shared memory where the block is located, see Figure 4 , specifically:
[0083] If the position of the current thread block is , the current block position is , the size of the thread block is .
[0084] This conversion ensures that each thread block correctly obtains the local data it is responsible for processing and saves it to the GPU high-speed shared memory of this block. This not only reduces the number of accesses to the global memory, but also improves the efficiency of data access, thereby accelerating the line search process.
[0085] S3-3, create a Boolean variable in the shared memory to control the scan loop of S3-4;
[0086] S3-4. All thread blocks continuously scan the grid area within their range and try to search for a connected line path, as follows:
[0087] The grid scanning process includes two stages: forward and reverse. Thread blocks in , in its corresponding A traversal search is performed within the range, first scanning forward (from left to right, from top to bottom), and then scanning backward (from right to left, from bottom to top). The scanning process is the same as the traditional generalized distance transform method.
[0088] It is worth noting that the scanning process is continuously executed in each thread block divided by S2 until the local Boolean variable shouldScan established in S3-3 is false, that is, the distance value in the same block no longer changes.
[0089] S3-5, detect whether the distance map data has changed during the scanning process, and if it has changed, copy the local distance map data in the shared memory to the global memory; specifically:
[0090] Before the grid scan begins, define a local Boolean variable hasChange that can only be accessed within this thread block. If the thread block modifies the distance data in the local shared memory during the scan, it will be assigned to true, otherwise it defaults to false. After the scan, if the hasChange variable is true, use the CAS (CompareAndSwap) operation to set the flag variable flag in the global memory to true. Since several thread blocks will perform this operation at the same time, this C operation ensures that as long as the hasChange variable in one thread block is true, the global flag variable flag is true. After the scan is completed, if flag is false, it means that all thread blocks in all blocks have not changed the distance map data.
[0091] S4: Determine whether the distance value in the distance map has changed. If it has changed, return to S3; otherwise, output the final distance map data. Specifically, it includes:
[0092] In the When executing step S3 for the second time, each block will obtain the latest distance map data of its adjacent blocks after the last round of scanning. , and then update the distance value of this block accordingly. For each block, the distance data used when performing grid scanning and searching for paths comes from , and by executing two rounds of S3, the shortest distance from the adjacent block to the target point can be propagated to the current block. In step S3 of the round, the generalized distance value of each block does not change. At this time, what is stored in the distance map is the shortest distance from all grid points to the target point (the starting point or the end point of the route).
[0093] S4-1. Update control array , Figure 5 (a) to Figure 5 (c) are schematic diagrams of inter-block distance propagation from the first round to the third round. After each round of S3, the control array Ctrl of each block will be updated. If the distance value of the block changes, the Ctrl value at the same position is true, then in the next round of search, the Ctrl values corresponding to the eight adjacent blocks around it are set to true.
[0094] S4-2. If the global marking variable flag is true, continue to execute the steps described in S3. If flag is false, the algorithm ends.
[0095] This embodiment also discloses a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the wide-area spatial railway line GPU parallel intelligent search method described above in this embodiment is implemented.
[0096] This embodiment also discloses an electronic device, including: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the wide-area spatial railway line GPU parallel intelligent search method as described above in this embodiment by executing the executable instructions.
[0097] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A GPU parallel intelligent search method for wide-area spatial railway lines, characterized in that: The steps include: S1. Establish a comprehensive geographic information model and initialize data; S2, divide the grid into levels; S3, parallelize the grid scanning in the GPU to update the distance map and find a feasible line path; S4, determine whether the distance value in the distance map has changed, if it has changed, return to S3, otherwise output the final distance map data; The S3 specifically includes the following steps: S3-1, checking whether the current block needs to update the distance map data according to the scan control array; S3-2, copy the local grid data and local distance map data corresponding to each block from the global memory to the shared memory where the block is located, and convert the global index into a local index during the copying; S3-3, create a Boolean variable in the shared memory to control the scan loop of S3-4; S3-4, all thread blocks continuously traverse and scan the grid within their range to find feasible line paths; S3-5, detecting whether the distance map data changes during the scanning process, and if so, copying the local distance map data in the shared memory to the global memory; The S4 includes: S4-1, update the scan control array; After each round of S3, the scan control array of each thread block will be updated; if the distance map data of the block in this round of search is changed, the scan control array value at the same position is true, then the scan control array values corresponding to the eight adjacent blocks around it will be set to true in the next round of search; S4-2, check whether the distance value of this round of scanning has been modified. If it has been modified, return to S3; if it has not been modified, the algorithm ends; The distance map data used by each block in the grid scanning search path comes from the distance map data of the previous round of scanning. By executing two rounds of S3 steps, the shortest distance to the target point in the adjacent block is propagated to the current block; after at most max{B r ,B c In the step of S3 of the}+1 round, the distance map data of each block will no longer change. At this time, what is stored in the distance map is the shortest distance from all grid points to the target point.
2. According to claim 1, a wide-area spatial railway line GPU parallel intelligent search method is characterized in that: The specific steps of S1 include: S1-1, establishing a comprehensive geographic information model array of grid data, wherein the grid data includes elevation, grid type, fill and cut volume, bridge and tunnel, land acquisition unit price and control factor information; S1-2, initialize the distance map array; S1-3, initialize the scan control array; S1-4, initialize the tag variable; The comprehensive geographic information model array is stored in the GPU global memory.
3. The wide-area spatial railway line GPU parallel intelligent search method according to claim 2 is characterized in that: The distance map data in the distance map array includes a generalized distance value and a relative row offset and a relative column offset of the next grid on the shortest path from the current grid to the target point.
4. The wide-area spatial railway line GPU parallel intelligent search method according to claim 2 is characterized in that: The grid level division of S2 includes the following steps: Divide the grid data in S1 into B r ×B c blocks; Each block is further subdivided into N×N thread blocks; Each thread block is responsible for processing a piece of grid data of size M×M; Based on the known size of the grid data, the block size, and the thread block size, the number of blocks contained in the grid data is calculated using the following formula: Where: B r Indicates the total number of rows of blocks into which the grid data is divided, B c Represents the total number of columns of the blocks into which the grid data is divided; NRow is the total number of rows of the grid data, NCol is the total number of columns of the grid data; M is the total number of rows or columns of the grid data contained in the thread block; N is the total number of rows or columns of the thread blocks contained in the block.
5. The wide-area spatial railway line GPU parallel intelligent search method according to claim 4, characterized in that: S3-1 specifically: Create a size B r ×B c The two-dimensional Boolean array of is used as the scan control array and stored in the global memory of the GPU; The scan control array is used to indicate whether the current block needs to be grid scanned to reduce the generalized distance value. The expression is as follows: Where: Ctrl(i,j) represents the scan control array; The scan control array optimizes the calculation process by marking which blocks need further processing. When the distance map data in a block changes, the Ctrl value of the block and its adjacent blocks will be set to true, triggering the next round of grid scanning and updating the distance map.
6. The wide-area spatial railway line GPU parallel intelligent search method according to claim 5, characterized in that: Before the grid scan described in S3-4 starts, a local Boolean variable that can only be accessed within the thread block is defined. If the thread block modifies the distance map data in the local shared memory during the scan, the variable is assigned a value of true, otherwise it defaults to false. The steps for detecting whether the distance map data has changed during the scanning process in S3-5 are as follows: after the scanning is completed, if the local Boolean variable is true, the mark variable in the global memory is set to true using the CAS operation; if the mark variable is false, it means that all thread blocks in all blocks have not changed the distance map data.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the wide-area spatial railway line GPU parallel intelligent search method described in any one of claims 1 to 6 is implemented.
8. An electronic device, characterized in that: include: processor; and a memory for storing executable instructions for the processor; Wherein, the processor is configured to execute the wide-area spatial railway line GPU parallel intelligent search method as described in any one of claims 1-6 by executing the executable instructions.
Citation Information
Patent Citations
Grid-free Galerkin method structural topology optimization method based on GPU parallel acceleration
CN103970960A
Railway line three-dimensional rapid modeling and dynamic updating method
CN114549717A