Parallel partitioning method and device for RDD structure post-stack data
By proposing a parallel blocking method in the RDD structure post-stack data processing, the problem of low efficiency in processing such data is solved, and efficient parallel blocking and subsequent three-dimensional processing is achieved.
Patent Information
- Application Number
- CN202311676070.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-07
- Publication Date
- 2025-06-10
AI Technical Summary
The traditional MPI parallel framework is inefficient when processing RDD structure post-stack data, especially in distributed storage and Π-Frame application systems.
A parallel chunking method for RDD structure stacked data is proposed. By obtaining the stacked data of RDD structure, determining the number and size of data blocks along the Inline and Crossline directions, calculating the partition number of each channel set, and generating a new
It realizes efficient parallel blocking of RDD structure data after stacking, improves processing efficiency and facilitates subsequent three-dimensional processing operations, such as three-dimensional space-time filtering.
Smart Images

Figure CN120123418A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of geophysical exploration, and more specifically, relates to a parallel block method and device for post-stack data with an RDD structure. Background Technique
[0002] With the development of geophysical exploration technology, three-dimensional processing of post-stack data has become a common means, such as three-dimensional spatio-temporal domain filtering. The popularization of three-dimensional processing technology has also brought the problem of a huge increase in the amount of calculation. Parallel processing is imperative, and the basis of parallel processing is the data block of post-stack data.
[0003] Post-stack data is usually a regular three-dimensional data volume. Usually, the data can be rectangularly blocked according to three variables: inline, crossline, and time (depth). There is usually an overlapping part between rectangular data blocks to prevent the problem of data mutation after processing. If it is a traditional MPI parallel framework, first, the partition number is looped, then the boundary points of the rectangular data block are determined according to the partition number, and then the corresponding data is read to complete the data block. However, for the characteristics of distributed storage and the Π-Frame application system (Sinopec), the post-stack data before processing is already a data with an RDD structure. If the implementation method of the traditional framework is still used, it will be very inefficient. Summary of the Invention
[0004] The object of the present invention is to propose a parallel block method and device for post-stack data with an RDD structure, and to realize the parallel block of post-stack data with an RDD structure efficiently.
[0005] To achieve the above object, in a first aspect, the present invention proposes a parallel block method for post-stack data with an RDD structure, including:
[0006] Obtain the post-stack data with an RDD structure, where the post-stack data is composed of <Key, Value> pairs, where Key is Inline and Crossline, and Value is composed of corresponding gathers;
[0007] Determine the number of data blocks, the size of each block, the size of the overlapping part between blocks, and the non-overlapping size of the last block along Inline and Crossline respectively;
[0008] Obtain the Inline and Crossline of a gather from the post-stack data, and calculate the partition number to which the gather belongs according to the size of each block and the size of the overlapping part between blocks. The partition number includes the block number of the gather in the Iinline direction and the block number in the Crossline direction;
[0009] Generate new <Key, Value> pairs, where the Key is the partition number of the trace gather and the Value is the trace gather;
[0010] Perform a partitioning operation based on the Key value of the trace gather to obtain a stacked trace gather with an RDD structure after partitioning.
[0011] Optionally, calculating the partition number to which the trace gather belongs includes:
[0012] Calculate the position information of the trace gather in the Iinline and Crossline directions respectively according to the size of each block and the size of the overlapping part between the blocks;
[0013] Determine the Iinline block number and Crossline block number of the trace gather according to the position information.
[0014] Optionally, calculating the position information of the trace gather in the Inline and Crossline directions respectively includes:
[0015] Calculate the block number of the trace gather in the Iinline direction or Crossline according to the size of each block and the size of the overlapping part between the blocks;
[0016] Calculate whether the block number is at the leftmost, middle, or rightmost end of the Inline or Crossline to obtain the first position information;
[0017] Calculate whether the position of the trace gather in the current block is in the left overlapping area, middle area, or right overlapping area to obtain the second position information.
[0018] Optionally, determining the Iinline block number and Crossline block number of the trace gather according to the position information includes:
[0019] Judge whether a new block number needs to be added according to the first position information and the second position information. If the second position information shows that the trace gather is in the left overlapping area of the current Iinline block or Crossline block, when the first position information shows that the block number is in the middle or rightmost end of the Inline or Crossline, add a new partition number in the Inline direction or Crossline direction.
[0020] Optionally, the method is executed based on distributed storage and the Spark framework.
[0021] In a second aspect, the present invention provides an electronic device, which includes:
[0022] At least one processor; and,
[0023] A memory communicatively connected to the at least one processor; wherein,
[0024] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the parallel block division method for post-stack data of the RDD structure according to any one of the first aspect.
[0025] In a third aspect, the present invention proposes a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the parallel block division method for post-stack data of the RDD structure according to any one of the first aspect.
[0026] In a fourth aspect, the present invention proposes a parallel block division device for post-stack data of the RDD structure, comprising:
[0027] A data acquisition module for acquiring post-stack data of the RDD structure, the post-stack data being composed of <Key, Value> pairs, where Key is Inline and Crossline, and Value is composed of corresponding gathers;
[0028] A parameter setting module for respectively determining the number of data blocks, the size of each block, the size of the overlapping part between blocks, and the non-overlapping size of the last block along Inline and Crossline;
[0029] A partition number calculation module for obtaining the Inline and Crossline of a gather from the post-stack data, and calculating the partition number to which the gather belongs according to the size of each block and the size of the overlapping part between blocks, the partition number including the block number of the gather in the Inline direction and the block number in the Crossline direction;
[0030] A generation module for generating a new <Key, Value> pair, where Key is the partition number of the gather and Value is the gather;
[0031] A partitioning module for performing a partitioning operation based on the Key value of the gather to obtain the post-stack gather of the RDD structure after block division.
[0032] Optionally, calculating the partition number to which the gather belongs includes:
[0033] Calculating the position information of the gather in Inline and Crossline respectively according to the size of each block and the size of the overlapping part between blocks;
[0034] Determining the Inline block number and Crossline block number of the gather according to the position information.
[0035] Optionally, calculate the position information of the trace gather in the Inline and Crossline directions respectively, including:
[0036] Calculate the block number of the trace gather in the Inline direction or Crossline according to the size of each block and the size of the overlapping part between blocks;
[0037] Calculate whether the block number is at the leftmost, middle or rightmost end of the Inline or Crossline to obtain the first position information;
[0038] Calculate whether the position of the trace gather in the current block is in the left overlapping area, middle area or right overlapping area to obtain the second position information.
[0039] The beneficial effects of the present invention are as follows:
[0040] The present invention realizes an efficient parallel block method for post-stack data in the RDD structure. After processing by this method, new post-stack data in the RDD will be obtained. The key of this data is the partition number, and the value is the trace gather, which can facilitate subsequent three-dimensional processing operations in parallel, such as three-dimensional spatio-temporal domain filtering.
[0041] The system of the present invention has other characteristics and advantages, which will be obvious from the accompanying drawings incorporated herein and the subsequent specific embodiments, or will be described in detail in the accompanying drawings incorporated herein and the subsequent specific embodiments. These accompanying drawings and specific embodiments are used together to explain the specific principles of the present invention. Description of the Drawings
[0042] By describing the exemplary embodiments of the present invention in more detail in conjunction with the accompanying drawings, the above and other objects, features and advantages of the present invention will become more obvious. In the exemplary embodiments of the present invention, the same reference numerals generally represent the same components.
[0043] Figure 1 A step diagram showing a parallel block method for post-stack data in the RDD structure according to the present invention is shown.
[0044] Figure 2 A schematic diagram of data block parameters in an embodiment of the present invention is shown.
[0045] Figure 3 A schematic diagram of the block number of a certain trace gather in an embodiment of the present invention is shown. Detailed Description of the Invention
[0046] The data in the RDD structure consists of <Key, Value> pairs. Here, the Key can be Inline or (Inline, Cossline), and the Value consists of the corresponding gathers. If the post-stack data in the RDD structure is chunked in the traditional way, each partition needs to obtain the corresponding gathers according to the data boundary points, which will involve a large number of addressing operations and is a very inefficient implementation method.
[0047] Based on the characteristics of the RDD structure itself, the present invention does not adopt the method of reading gathers by looping through partitions. Instead, it performs concurrent operations on the gathers to determine which partitions the gathers belong to, then labels the gathers with partition labels, and completes data chunking through data grouping operations.
[0048] The present invention will be described in more detail below with reference to the accompanying drawings. Although the preferred embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided to make the present invention more thorough and complete, and to fully convey the scope of the present invention to those skilled in the art.
[0049] Embodiment 1
[0050] As Figure 1 shown, this embodiment provides a parallel chunking method for post-stack data in the RDD structure, including:
[0051] S1: Obtain the post-stack data in the RDD structure. The post-stack data consists of <Key, Value> pairs, where the Key is Inline and Crossline, and the Value consists of the corresponding gathers;
[0052] S2: Determine the number of data chunks, the size of each chunk, the size of the overlapping part between chunks, and the non-overlapping size of the last chunk along Inline and Crossline respectively;
[0053] S3: Obtain the Inline and Crossline of a gather from the post-stack data, and calculate the partition numbers to which the gather belongs according to the size of each chunk and the size of the overlapping part between chunks. The partition numbers include the chunk numbers of the gather in the Inline direction and in the Crossline direction;
[0054] In this step, calculating the partition numbers to which the gather belongs includes:
[0055] Calculating the position information of the gather in Inline and Crossline respectively according to the size of each chunk and the size of the overlapping part between chunks;
[0056] Determine the Iinline block number and Crossline block number of the trace gather according to the position information.
[0057] Among them, calculate the position information of the trace gather in Inline and Crossline respectively, including:
[0058] According to the size of each block and the size of the overlapping part between blocks, calculate the block number of the trace gather in the Inline direction or Crossline;
[0059] Calculate whether the block number is at the leftmost, middle or rightmost of Inline or Crossline to obtain the first position information;
[0060] Calculate whether the position of the trace gather in the current block is the left overlapping area, the middle area or the right overlapping area to obtain the second position information.
[0061] Determining the Iinline block number and Crossline block number of the trace gather according to the position information includes:
[0062] Judge whether a new block number needs to be added according to the first position information and the second position information. If the second position information shows that the trace gather is in the left overlapping area of the current Inline block or Crossline block, then when the first position information shows that the block number is in the middle or rightmost of Inline or Crossline, add a new partition number in the Inline direction or Crossline direction.
[0063] S4: Generate a new <Key, Value> pair, where Key is the partition number of the trace gather and Value is the trace gather;
[0064] S5: Perform a partitioning operation based on the Key value of the trace gather to obtain a partitioned RDD structure stacked post-stack trace gather.
[0065] Preferably, the method of this embodiment is executed based on distributed storage and the Spark framework.
[0066] Embodiment 2
[0067] This embodiment provides a parallel chunking method for post-stack data with an RDD structure. The implementation of this method is based on distributed storage and the Spark framework and is implemented in Scala. The implementation of this method does not involve the generation process of post-stack data with an RDD structure, but rather the subsequent chunking operation on post-stack data with an RDD structure. The post-stack data with an RDD structure processed by this method consists of <Key, Value> pairs. The Key can be Inline or (Inline, Crossline), and the Value consists of the corresponding gathers. This method only chunks the data in the Inline and Crossline directions. The chunking in the time (depth) direction is no different from the conventional method (because a complete gather can be obtained each time). Therefore, the partition numbers of the data chunks (Blocks) will be composed of (Block_Inline, Block_Corssline).
[0068] This method specifically includes the following steps:
[0069] (1) First, determine the number of data chunks (Blocks), the size (Size) of each chunk, the size of the overlapping part (Size_Shadow), and the non-overlapping size of the last chunk (Size_Last), and count them separately along Inline and Crossline, as Figure 2 shown. The shaded part in the figure is the overlapping part of the data blocks.
[0070] (2) Second, perform flatMap (a keyword in Scala for parallel operations) on the post-stack data to obtain the (Inline, Crossline) of each gather.
[0071] Next, three values need to be calculated. Taking Inline as an example, the first is to calculate the Block_Inline number of this gather in the Inline direction according to (Size - Size_Shadow). The second is to calculate whether this Block_Inline is at the leftmost, middle, or rightmost in Inline (i.e., Block_Location1). The third is to calculate whether the position of this gather in the current Block_Inline is in the left shaded area, middle area, or right shaded area (i.e., Block_Location2). Similar operations are performed in the Crossline direction.
[0072] (3) Jointly judge whether a new partition number needs to be added based on the two values of Block_Loacation1 and Block_Location2. If Block_Location2 shows that it is in the left overlapping area of the current Block, then when Block_Location1 is in the middle and rightmost, a new partition number (Block - 1) needs to be added.
[0073] (4) Finally, the partition number to which the trace gather belongs will be composed of (Block_Inline, Block_Crossline). As shown below, both Block_Inline and Block_Crossline can be either 1 or 2. Therefore, there are three possibilities for the partition number to which the trace gather belongs: 1, 2, and 4. Figure 3 As shown, both Block_Inline and Block_Crossline can be either 1 or 2. Therefore, there are three possibilities for the partition number to which the trace gather belongs: 1, 2, and 4.
[0074] (5) Generate new <Key, Value> pairs, where Key is (Block_Inline, Block_Crossline) and Value is the trace gather. Finally, perform the groupByKey (grouping based on the Key value) operation to obtain the post-stack trace gather in the RDD structure after partitioning. This RDD data can be conveniently used for the next three-dimensional data processing.
[0075] Embodiment 3
[0076] This embodiment provides a parallel partitioning device for post-stack data in RDD structure, including:
[0077] A data acquisition module, configured to acquire post-stack data in RDD structure, where the post-stack data is composed of <Key, Value> pairs, with Key being Inline and Crossline, and Value being composed of corresponding trace gathers;
[0078] A parameter setting module, configured to respectively determine the number of data partitions, the size of each partition, the overlapping size between partitions, and the non-overlapping size of the last partition along Inline and Crossline;
[0079] A partition number calculation module, configured to obtain the Inline and Crossline of a trace gather from the post-stack data, and calculate the partition number to which the trace gather belongs according to the size of each partition and the overlapping size between partitions. The partition number includes the partition number of the trace gather in the Inline direction and the partition number in the Crossline direction;
[0080] A generation module, configured to generate new <Key, Value> pairs, where Key is the partition number of the trace gather and Value is the trace gather;
[0081] A partitioning module, configured to perform a partitioning operation based on the Key value of the trace gather to obtain the post-stack trace gather in the RDD structure after partitioning.
[0082] In this embodiment, calculating the partition number to which the trace gather belongs includes:
[0083] Calculating the position information of the trace gather in Inline and Crossline respectively according to the size of each partition and the overlapping size between partitions;
[0084] Determine the Iinline block number and Crossline block number of the trace gather according to the position information.
[0085] In this embodiment, calculating the position information of the trace gather in Inline and Crossline respectively includes:
[0086] Calculate the block number of the trace gather in the Inline direction or Crossline according to the size of each block and the size of the overlapping part between blocks;
[0087] Calculate whether the block number is at the leftmost, middle or rightmost of Inline or Crossline to obtain the first position information;
[0088] Calculate whether the position of the trace gather in the current block is the left overlapping area, the middle area or the right overlapping area to obtain the second position information.
[0089] In this embodiment, the determining the Iinline block number and Crossline block number of the trace gather according to the position information includes:
[0090] Judge whether a new block number needs to be added according to the first position information and the second position information. If the second position information shows that the trace gather is in the left overlapping area of the current Inline block or Crossline block, when the first position information shows that the block number is in the middle or rightmost of Inline or Crossline, add a new partition number in the Inline direction or Crossline direction.
[0091] In this embodiment, the device is implemented by using distributed storage and the Spark framework.
[0092] Embodiment 4
[0093] This embodiment provides an electronic device, and the electronic device includes:
[0094] At least one processor; and,
[0095] A memory communicatively connected to the at least one processor; wherein,
[0096] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the parallel block method of the post-stack data of the RDD structure described in Embodiment 1 or 2.
[0097] An electronic device according to an embodiment of the present disclosure includes a memory and a processor. The memory is used to store non-transitory computer-readable instructions. Specifically, the memory may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.
[0098] The processor may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions. In an embodiment of the present disclosure, the processor is used to run the computer-readable instructions stored in the memory.
[0099] Those skilled in the art should understand that, in order to solve the technical problem of how to obtain good user experience effects, well-known structures such as communication buses and interfaces may also be included in this embodiment, and these well-known structures should also be included in the protection scope of the present disclosure.
[0100] For a detailed description of this embodiment, reference may be made to the corresponding descriptions in the foregoing embodiments, and details will not be repeated here.
[0101] Embodiment 5
[0102] This embodiment provides a non-transitory computer-readable storage medium that stores computer instructions for causing a computer to execute the parallel block division method for post-stack data of the RDD structure described in Embodiment 1 or 2.
[0103] A computer-readable storage medium according to an embodiment of the present disclosure stores non-transitory computer-readable instructions thereon. When the non-transitory computer-readable instructions are run by a processor, all or part of the steps of the methods of the foregoing embodiments of the present disclosure are executed.
[0104] The above-mentioned computer-readable storage media include, but are not limited to: optical storage media (such as CD-ROM and DVD), magneto-optical storage media (such as MO), magnetic storage media (such as magnetic tapes or external hard drives), media with built-in rewritable non-volatile memory (such as memory cards), and media with built-in ROM (such as ROM cartridges).
[0105] The embodiments of the present invention have been described above. The above description is exemplary and not exhaustive, and is also not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments.
Claims
1. A parallel block method for post-stack data of RDD structure, characterized in that, it includes: Obtain the post-stack data of the RDD structure, where the post-stack data consists of <Key, Value> pairs, where Key is Inline and Crossline, and Value is composed of corresponding gathers; Determine the number of data blocks, the size of each block, the size of the overlapping part between blocks, and the non-overlapping size of the last block along Inline and Crossline respectively; Obtain the Inline and Crossline of a gather from the post-stack data, and calculate the partition number to which the gather belongs according to the size of each block and the size of the overlapping part between blocks. The partition number includes the block number of the gather in the Inline direction and the block number in the Crossline direction; Generate a new <Key, Value> pair, where Key is the partition number of the gather, and Value is the gather; Perform a partition operation based on the Key value of the gather to obtain the post-stack gathers of the RDD structure after block division.
2. The parallel block method for post-stack data of the RDD structure according to claim 1, characterized in that, The calculation of the partition number to which the gather belongs includes: Calculate the position information of the gather in Inline and Crossline respectively according to the size of each block and the size of the overlapping part between blocks; Determine the Inline block number and Crossline block number of the gather according to the position information.
3. The parallel block method for post-stack data of the RDD structure according to claim 2, characterized in that, Calculating the position information of the gather in Inline and Crossline respectively includes: Calculate the block number of the gather in the Inline direction or Crossline according to the size of each block and the size of the overlapping part between blocks; Calculate whether the block number is at the leftmost, middle or rightmost end of Inline or Crossline to obtain the first position information; Calculate whether the position of the gather in the current block is in the left overlapping area, middle area or right overlapping area to obtain the second position information.
4. The parallel block method for post-stack data of the RDD structure according to claim 3, characterized in that, The determination of the Inline block number and Crossline block number of the gather according to the position information includes: Judge whether a new block number needs to be added according to the first position information and the second position information. If the second position information shows that the gather is in the left overlapping area of the current Inline block or Crossline block, then when the first position information shows that the block number is in the middle or rightmost end of Inline or Crossline, add a new partition number in the Inline direction or Crossline direction.
5. The parallel block method for post-stack data of the RDD structure according to any one of claims 1-4, characterized in that, The method is executed based on distributed storage and the Spark framework.
6. An electronic device, It is characterized in that the electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the parallel block method for post-stack data of the RDD structure according to any one of claims 1-5.
7. A non-transitory computer-readable storage medium It is characterized in that the non-transitory computer-readable storage medium stores computer instructions for causing a computer to execute the parallel block method for post-stack data of the RDD structure according to any one of claims 1-5.
8. A parallel block device for post-stack data of the RDD structure It is characterized in that it includes: a data acquisition module for acquiring post-stack data of the RDD structure, the post-stack data being composed of <Key, Value> pairs, where Key is Inline and Crossline, and Value is composed of corresponding gathers; a parameter setting module for respectively determining the number of data blocks, the size of each block, the size of the overlapping part between blocks, and the non-overlapping size of the last block along Inline and Crossline; a partition number calculation module for obtaining the Inline and Crossline of a gather from the post-stack data, and calculating the partition number to which the gather belongs according to the size of each block and the size of the overlapping part between blocks, the partition number including the block number of the gather in the Inline direction and the block number in the Crossline direction; a generation module for generating a new <Key, Value> pair, where Key is the partition number of the gather and Value is the gather; a partition module for performing a partitioning operation based on the Key value of the gather to obtain the post-stack gathers of the RDD structure after being blocked.
9. The parallel block device for post-stack data of the RDD structure according to claim 8 It is characterized in that calculating the partition number to which the gather belongs includes: calculating the position information of the gather in Inline and Crossline respectively according to the size of each block and the size of the overlapping part between blocks; determining the Inline block number and Crossline block number of the gather according to the position information.
10. The parallel block device for post-stack data of the RDD structure according to claim 9 It is characterized in that calculating the position information of the gather in Inline and Crossline respectively includes: calculating the block number of the gather in the Inline direction or Crossline according to the size of each block and the size of the overlapping part between blocks; calculating whether the block number is at the leftmost, middle or rightmost of Inline or Crossline to obtain the first position information; calculating whether the position of the gather in the current block is in the left overlapping area, middle area or right overlapping area to obtain the second position information.
Citation Information
Cited By
Seismic data distributed partition Halo region overlapping optimization method and device
CN121049973A