Seismic data parallel circulation method and device based on RDD

Through the parallel flow method based on RDD, the track set data structure of seismic data is constructed, and the parallel data flow of RDD is used to solve the problem of frequent reading and writing of data and parallel processing in the prior art, and efficient storage and computing efficiency is achieved.

CN120123311APending Publication Date: 2025-06-10CHINA PETROLEUM & CHEMICAL CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311684654.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-08
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

In the existing seismic data processing methods, frequent read and write operations of temporary files lead to large storage requirements and low processing efficiency, and parallel processing is difficult, making it difficult to meet the needs of large-scale seismic data processing.

Method used

By using the RDD-based parallel flow method, by constructing an RDD data model, seismic data is organized according to a track set, and the data transfer between multiple processing modules is completed by using the parallel data flow of RDD, avoiding frequent data read and write operations, and parallelized large-scale data processing is realized.

Benefits of technology

This method saves storage space, improves computing efficiency, simplifies multi-work management, and realizes the need for parallelized large-scale seismic data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123311A_ABST
    Figure CN120123311A_ABST
Patent Text Reader

Abstract

The invention relates to the field of seismic exploration data processing, and particularly discloses a seismic data parallel circulation method and device based on RDD, and the method comprises the steps: constructing an RDD seismic data set structure, and recording the information of a plurality of gathers; constructing a gather data structure to record a plurality of pieces of seismic trace number information and seismic data file headers in each gather; and constructing an RDD transfer function and transferring data among the plurality of processing modules in parallel. According to the seismic data parallel circulation method based on the RDD provided by the invention, the seismic data is utilized to construct the RDD data model according to the trace gather, and the data transmission among a plurality of processing modules is completed by utilizing the parallel data circulation of the RDD data model, so that the frequent data read-write operation in a complex processing flow is avoided, and meanwhile, the parallel large-scale data processing requirement is realized; the storage space is saved, and the calculation efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of seismic exploration data processing, and in particular to a RDD-based seismic data parallel flow method and device. Background Art

[0002] The seismic data processing process is complex and requires multiple processing methods and processing modules. In order to improve the processing efficiency and processing effect, multiple processing modules are usually connected in series into a processing flow to complete the entire processing. To realize the processing of complex processing flows, the core is the data flow between processing modules. The current commonly used method is to output the processing results of each module in the processing flow to the storage disk to form a temporary file, and the next module reads the temporary file for subsequent processing, and the data flow in the processing flow is realized through the temporary file. This method requires a large temporary storage space, and at the same time, because each module requires frequent read and write operations, it affects the processing efficiency. In order to solve the temporary file problem, another commonly used method is to generate a unified execution program by recompiling multiple modules in the processing flow. Data flow is completed through the memory in the program. This processing method reduces the read and write operations of temporary files, has low storage requirements, and high computational efficiency. However, parallel processing is difficult. For large-scale seismic data processing, the data needs to be divided into a large number of small files, and a processing flow is started for each file for processing. The operation is complicated and multi-job management is difficult.

[0003] The abstract Resiliennt Distributed Datasets (RDD) in Spark parallel technology is an efficient parallel data management method. It provides data management and flow methods for large-scale data parallel processing and is widely used in big data processing.

[0004] Based on this technical background, the present invention studies a RDD-based seismic data parallel flow method and device. Summary of the invention

[0005] In view of the shortcomings of the prior art, the present invention provides a method and device for parallel flow of seismic data based on RDD. The method uses seismic data to construct an RDD data model according to the track set, and uses the parallel data flow of the RDD data model to complete data transfer between multiple processing modules, thereby avoiding frequent reading and writing operations of intermediate data in complex processing flows, and at the same time realizing parallel large-scale data processing requirements, saving storage space and improving computing efficiency.

[0006] In order to achieve the above object, a first aspect of the present invention provides a RDD-based seismic data parallel flow method, comprising:

[0007] Construct an RDD seismic data set structure to record multiple gather information;

[0008] Construct a gather data structure to record multiple seismic channel information and seismic data file headers in each gather;

[0009] Construct RDD transfer functions to transfer data between multiple processing modules in parallel.

[0010] A second aspect of the present invention provides a seismic data parallel flow device based on RDD, comprising:

[0011] The seismic data set structure building module is used to build the RDD seismic data set structure to record multiple gather information;

[0012] A gather data structure construction module, used to construct a gather data structure to record multiple seismic channel information and seismic data file headers in each gather;

[0013] The data flow module is used to build an RDD transfer function to transfer data between multiple processing modules in parallel.

[0014] A third aspect of the present invention provides an electronic device, the electronic device comprising:

[0015] A memory storing executable instructions;

[0016] A processor runs the executable instructions in the memory to implement the RDD-based parallel flow method for seismic data described in the first aspect.

[0017] A fourth aspect of the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the RDD-based parallel flow method for seismic data described in the first aspect.

[0018] The beneficial effects of the present invention include:

[0019] (1) The RDD-based seismic data parallel flow method proposed in the present invention uses seismic data to construct an RDD data model according to the track set, and uses the parallel data flow of the RDD data model to complete the data transfer between multiple processing modules, avoiding the frequent reading and writing operations of intermediate data in complex processing flows, while realizing the parallel large-scale data processing requirements, saving storage space, and improving computing efficiency.

[0020] (2) The RDD-based parallel flow method for seismic data proposed in the present invention adds file header information to the seismic data gather record design in response to the requirements of parallel computing processing, so that each RDD gather record has file header data, providing necessary global information for parallel seismic data processing.

[0021] (3) In the RDD-based seismic data parallel flow method proposed in the present invention, the file header is stored in the form of a byte array, which is convenient for reading and transmitting data. The integrity of the RDD-based seismic data flow is ensured through the parallel setting of the file header.

[0022] (4) The RDD-based parallel flow method for seismic data proposed in the present invention designs a data structure for storing the trace header and trace data separately, in view of the different structures of the trace header and the trace data body. The trace header data is stored in the form of a byte array, and the trace data body is stored in the form of a floating-point array, thereby improving the efficiency of parallel flow of seismic data.

[0023] Other features and advantages of the present invention will be described in detail in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The above and other objects, features and advantages of the present invention will become more apparent through a more detailed description of exemplary embodiments of the present invention with reference to the accompanying drawings.

[0025] Figure 1 This is a flow chart of the RDD-based parallel flow method for seismic data proposed in the present invention. DETAILED DESCRIPTION

[0026] The preferred embodiments of the present invention will be described in more detail below. Although the preferred embodiments of the present invention are described below, it should be understood that the present invention can be implemented in various forms and should not be limited to the embodiments set forth herein.

[0027] The present invention provides a parallel flow method of seismic data based on RDD, such as Figure 1 ,include:

[0028] Construct an RDD seismic data set structure to record multiple gather information;

[0029] Construct a gather data structure to record multiple seismic channel information and seismic data file headers in each gather;

[0030] Construct RDD transfer functions to transfer data between multiple processing modules in parallel.

[0031] In the present invention, seismic data is used to construct an RDD data model according to the track gather, and the parallel data flow of the RDD data model is used to complete the data transfer between multiple processing modules, avoiding frequent reading and writing operations of intermediate data in complex processing flows, while realizing parallel large-scale data processing requirements, saving storage space, and improving computing efficiency.

[0032] According to the present invention, each track gather is composed of a plurality of seismic track data with the same order;

[0033] Each seismic trace data is a seismic wave signal received by a detector over a period of time during seismic exploration.

[0034] In the present invention, the basic unit of seismic data is a trace, and each trace represents a seismic wave signal received by a detector for a period of time in seismic exploration. In seismic data processing, seismic trace data with the same order are usually combined into a trace set, which is the most commonly used data unit in seismic data processing.

[0035] According to the present invention, each gather information includes a unique identification number and a gather data body;

[0036] Each gather information record adopts the key-value pair method, with the key recording the unique identification number of the gather and the value recording the gather data body of the gather.

[0037] According to the present invention, each seismic trace data information includes a trace header and a trace data body;

[0038] The trace header is a data structure with a fixed byte length, and the trace data is a one-dimensional floating point array;

[0039] Each seismic channel number information is recorded using the Record class;

[0040] The trace header is stored in the form of a byte array, and the trace data body is stored in the form of a floating point array.

[0041] The method proposed in the present invention designs a track gather data structure in which the track header and the track data are stored separately, in view of the different structures of the track header and the track data body. The track header data is stored in the form of a byte array, and the track gather data body is stored in the form of a floating point array, thereby improving the efficiency of parallel flow of seismic data.

[0042] Preferably, the trace header records the acquisition information of the trace data body, and the acquisition information includes coordinate position, sampling interval, sampling length and grid information;

[0043] Channel data records the amplitude value of the sampling point;

[0044] The Record class also records the number of channels in each channel set and the sample data of each seismic channel.

[0045] According to the present invention, the seismic data file header is a file header that describes the entire seismic data information;

[0046] The earthquake data file header is recorded by the Record class;

[0047] The header of the seismic data file is stored in a byte array and is set up in parallel.

[0048] The method proposed in the present invention adds file header information to the seismic data gather record design in response to the requirements of parallel computing processing, so that each gather record of RDD has file header data, providing necessary global information for parallel seismic data processing.

[0049] In the present invention, the key value of each record in the RDD data set is determined by the track set identification number, and the value records a track set data body, including multiple track data, and a data storage structure needs to be reasonably designed. Each track of seismic data includes a track header and a track data body. The track header records the acquisition information of the track data, including coordinate position, sampling interval, sampling length, grid information and other information. A track set record class Record is constructed. The Record class records the seismic data track header and track data, and records the basic information such as the number of tracks of the track set data and the sample point data of each seismic data.

[0050] In the present invention, the basic unit of the seismic data structure is the track, and each seismic data file has a file header that describes the entire seismic data information. When the seismic data is processed in parallel, the global information of the seismic data may be needed in the calculation process of each track set, which needs to be obtained from the file header. In view of the requirements of parallel computing processing, file header information is added to the seismic data track set record Record design to implement RDD. Each track set record has file header data, providing necessary global information for parallel seismic data processing.

[0051] In the method proposed in the present invention, the file header is stored in the form of a byte array, which is convenient for reading and transmitting data. The integrity of the seismic data flow based on RDD is ensured through the parallel setting of the file header.

[0052] Preferably, the RDD transfer function includes:

[0053] getInputRdd() function is used to obtain the data passed by the previous module;

[0054] The setOutputRdd() function is used to pass the data calculated by this module to the next module.

[0055] In the present invention, the seismic data processing flow includes multiple processing modules, each processing module is an independent program, and after the RDD based on the seismic data is constructed, it is necessary to complete the transfer of RDD between multiple modules. The present invention designs a basic class that supports the development of seismic data processing modules, and designs RDD transfer functions in the module basic class: getInputRdd() and setOutputRdd(). All processing modules are developed based on the basic class, so that RDD transfer between modules in complex processing flows can be realized.

[0056] The present invention will be described in more detail below by way of examples.

[0057] Embodiment 1:

[0058] This embodiment provides 一 种 The specific steps of the parallel flow method of seismic data based on RDD are as follows:

[0059] The first step is to determine the seismic gather data: there are many types of seismic data gathers, and determining the gather type is the prerequisite for building RDD; taking the common center point gather type of seismic data as an example, the common center point gather type is determined by two gather identifiers, the horizontal inline number and the vertical xline number;

[0060] The second step is to construct the track set type Record: extract the inline number and xline number of each track set, construct the key value of RDD, extract all seismic track data with the same inline number and xline number as the seismic data of a track set; split the track header and track data body of the seismic data, convert the track header structure into a byte array, and load the track header byte array of all tracks in the track set into the track header array in the Record class; load all track data in the track set into the track data array in the Record class; after completing the loading of the track header data and track data, set the number of seismic data tracks, the number of sampling points, the sampling interval and other information to complete the construction of the Record seismic data;

[0061] Step 3: File header setting: extract the seismic data file header, convert it into byte data, and load it into the Record of each seismic data RDD to complete the file header data loading;

[0062] The fourth step is RDD transfer: for a specific module, if it is the initial module of the processing flow, after completing the RDD construction, new RDD data is generated after processing by this module and passed to the next module through setOutputRdd(); if it is an intermediate module in the processing flow, the RDD data of the previous module is obtained through the getInputRdd() function, and after completing the processing of this module, it is also passed to the next module through setOutputRdd(); through the transfer of RDD, the data flow of the entire processing flow is realized, ensuring the processing of complex processing flows.

[0063] Embodiment 2:

[0064] like Figure 1 As shown, this embodiment provides a parallel flow method of seismic data based on RDD, including:

[0065] Construct an RDD seismic data set structure to record multiple gather information;

[0066] Construct a gather data structure to record multiple seismic channel information and seismic data file headers in each gather;

[0067] Build RDD transfer functions to transfer data between multiple processing modules in parallel;

[0068] Each gather consists of multiple seismic trace data with the same order;

[0069] Each seismic channel data is a seismic wave signal received by a detector over a period of time during seismic exploration;

[0070] Each gather information includes a unique identification number and a gather data body;

[0071] Each gather information record adopts the key-value pair method, with the key recording the unique identification number of the gather and the value recording the gather data body of the gather;

[0072] Each seismic channel data information includes a channel header and a channel data body;

[0073] The trace header is a data structure with a fixed byte length, and the trace data is a one-dimensional floating point array;

[0074] Each seismic channel number information is recorded using the Record class;

[0075] The trace header is stored in the form of a byte array, and the trace data body is stored in the form of a floating point array;

[0076] The trace header records the acquisition information of the trace data body, which includes coordinate position, sampling interval, sampling length and grid information;

[0077] Channel data records the amplitude value of the sampling point;

[0078] The Record class also records the number of channels in each channel set and the sample data of each seismic channel;

[0079] The seismic data file header is a file header that describes the entire seismic data information;

[0080] The earthquake data file header is recorded by the Record class;

[0081] The header of the seismic data file is stored in a byte array and is set up in parallel;

[0082] RDD transfer functions include:

[0083] getInputRdd() function is used to obtain the data passed by the previous module;

[0084] The setOutputRdd() function is used to pass the data calculated by this module to the next module.

[0085] Embodiment three:

[0086] This embodiment provides a seismic data parallel flow device based on RDD, including:

[0087] The seismic data set structure building module is used to build the RDD seismic data set structure to record multiple gather information;

[0088] A gather data structure construction module, used to construct a gather data structure to record multiple seismic channel information and seismic data file headers in each gather;

[0089] The data flow module is used to build RDD transfer functions to flow data between multiple processing modules in parallel;

[0090] Each gather consists of multiple seismic trace data with the same order;

[0091] Each seismic channel data is a seismic wave signal received by a detector over a period of time during seismic exploration;

[0092] Each gather information includes a unique identification number and a gather data body;

[0093] Each gather information record adopts the key-value pair method, with the key recording the unique identification number of the gather and the value recording the gather data body of the gather;

[0094] Each seismic channel data information includes a channel header and a channel data body;

[0095] The trace header is a data structure with a fixed byte length, and the trace data is a one-dimensional floating point array;

[0096] Each seismic channel number information is recorded using the Record class;

[0097] The trace header is stored in the form of a byte array, and the trace data body is stored in the form of a floating point array;

[0098] The trace header records the acquisition information of the trace data body, which includes coordinate position, sampling interval, sampling length and grid information;

[0099] Channel data records the amplitude value of the sampling point;

[0100] The Record class also records the number of channels in each channel set and the sample data of each seismic channel;

[0101] The seismic data file header is a file header that describes the entire seismic data information;

[0102] The earthquake data file header is recorded by the Record class;

[0103] The header of the seismic data file is stored in a byte array and is set up in parallel;

[0104] RDD transfer functions include:

[0105] getInputRdd() function is used to obtain the data passed by the previous module;

[0106] The setOutputRdd() function is used to pass the data calculated by this module to the next module.

[0107] Embodiment 4:

[0108] An embodiment of the present invention provides an electronic device including a memory and a processor.

[0109] A memory storing executable instructions;

[0110] The processor runs the executable instructions in the memory to implement the RDD-based seismic data parallel flow method.

[0111] The memory is used to store non-temporary computer-readable instructions. Specifically, the memory may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may, for example, include random access memory (RAM) and / or cache memory (cache), etc. The non-volatile memory may, for example, include read-only memory (ROM), hard disk, flash memory, etc.

[0112] The processor may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions. In one embodiment of the present invention, the processor is used to run the computer-readable instructions stored in the memory.

[0113] Those skilled in the art should be able to understand that in order to solve the technical problem of how to obtain a good user experience, the present embodiment may also include well-known structures such as a communication bus and an interface, and these well-known structures should also be included in the protection scope of the present invention.

[0114] For detailed description of this embodiment, reference may be made to the corresponding descriptions in the aforementioned embodiments, which will not be repeated here.

[0115] Embodiment five:

[0116] An embodiment of the present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, a parallel flow method of seismic data based on RDD is implemented.

[0117] The computer-readable storage medium according to the embodiment of the present invention stores non-transitory computer-readable instructions, and when the non-transitory computer-readable instructions are executed by a processor, all or part of the steps of the above-mentioned methods of the embodiments of the present invention are executed.

[0118] The above-mentioned computer-readable storage media include, but are not limited to: optical storage media (e.g., CD-ROM and DVD), magneto-optical storage media (e.g., MO), magnetic storage media (e.g., magnetic tape or mobile hard disk), media with built-in rewritable non-volatile memory (e.g., memory card) and media with built-in ROM (e.g., ROM box).

[0119] The embodiment of the present invention proposes an RDD-based seismic data parallel flow method, which utilizes seismic data to construct an RDD data model according to the track set, and utilizes the parallel data flow of the RDD data model to complete data transfer between multiple processing modules, thereby avoiding frequent reading and writing operations of intermediate data in complex processing flows, and at the same time realizing parallel large-scale data processing requirements, saving storage space, and improving computing efficiency.

[0120] The embodiments of the present invention have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments.

Claims

1. A parallel transfer method for seismic data based on RDD, characterized in that, it includes: Construct an RDD seismic data set structure to record multiple gather information; Construct a gather data structure to record the number of multiple seismic traces and the seismic data file header in each gather; Construct an RDD transfer function to transfer data between multiple processing modules in parallel.

2. The method according to claim 1, characterized in that, each gather is composed of multiple seismic trace data with the same sorting; each seismic trace data is a seismic wave signal received by a geophone for a period of time during seismic exploration.

3. The method according to claim 1, characterized in that, each gather information includes a unique identification number and a gather data body; each gather information record adopts the key-value pair method, uses the key to record the unique identification number of the gather, and uses the value to record the gather data body of the gather.

4. The method according to claim 3, characterized in that, each seismic trace data information includes a trace header and a trace data body; the trace header is a data structure with a fixed byte length, and the trace data is a one-dimensional floating-point array; each seismic trace number information is recorded using the Record class; the trace header is stored in the form of a byte array, and the trace data body is stored in the form of a floating-point array.

5. The method according to claim 4, characterized in that, the trace header records the acquisition information of the trace data body, and the acquisition information includes coordinate position, sampling interval, sampling length, and grid information; the trace data records the amplitude value of the sampling point; the Record class also records the number of traces in each gather and the sample point data of each seismic trace.

6. The method according to claim 5, characterized in that, the seismic data file header is a file header describing the entire seismic data information; the seismic data file header is recorded by the Record class; the seismic data file header is stored in the form of a byte array and is set in parallel.

7. The method according to claim 6, characterized in that, the RDD transfer function includes: The getInputRdd() function is used to obtain the data transferred from the previous module; The setOutputRdd() function is used to transfer the calculated data of this module to the next module.

8. A parallel transfer device for seismic data based on RDD, characterized in that, it includes: A seismic data set structure construction module for constructing an RDD seismic data set structure to record multiple gather information; A gather data structure construction module for constructing a gather data structure to record the number of multiple seismic traces and the seismic data file header in each gather; A data transfer module for constructing an RDD transfer function to transfer data between multiple processing modules in parallel.

9. An electronic device, characterized in that, the electronic device includes: A memory storing executable instructions; A processor that runs the executable instructions in the memory to implement the parallel transfer method for seismic data based on RDD according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which when executed by a processor implements the RDD-based parallel transfer method of seismic data according to any one of claims 1-7.