PySpark-based seismic data batch processing method and device

By adopting the PySpark-based seismic data batch processing method in seismic data processing, segy data is processed in blocks and parallel, solving the problem of low processing efficiency of massive seismic data and achieving efficient calculation and data processing.

CN120122173APending Publication Date: 2025-06-10CHINA PETROLEUM & CHEMICAL CORP +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202311686716.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-08
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently process and analyze massive seismic data, especially in multi-seismic acquisition and high-density acquisition environments, where the calculation efficiency and processing speed are insufficient.

Method used

Using the seismic data batch processing method based on PySpark, segy data is stored in blocks into different RDDs, core functions of data processing are defined, and module development templates are built to realize parallel processing of the computing process, and Spark's cluster computing capabilities are utilized.

Benefits of technology

Through parallel processing and distributed computing, the calculation efficiency and processing speed of seismic data processing are significantly improved, and the development efficiency and data processing efficiency are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120122173A_ABST
    Figure CN120122173A_ABST
Patent Text Reader

Abstract

The invention relates to the field of seismic exploration data processing, and particularly discloses a PySpark-based seismic data batch processing method and device, and the method comprises the steps: loading segment data through employing a segment library; the segment data is stored in different RDDs in a partitioning mode; defining, by a developer, a core function of data processing performed by each RDD block; constructing a module development template based on a core function of data processing performed by each RDD block; and outputting and storing a processing result of the module development template. According to the method provided by the invention, the parallelization processing of the calculation process is realized by carrying out block storage on the segment data, defining a core function for data processing and constructing a module development template, and meanwhile, the calculation efficiency and the processing speed are improved by utilizing the cluster calculation capability of Spark.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of seismic exploration data processing, and particularly to a method and device for batch processing of seismic data based on PySpark. Background Art

[0002] Seismic data processing involves a large number of computational operations, including matrix operations, Fourier transforms, convolutions, etc., for implementing algorithms such as preprocessing, denoising, inversion, imaging, etc. To implement these algorithms, the programming language should have efficient numerical computing capabilities and be able to execute these computational operations quickly to improve the processing speed and efficiency. Moreover, seismic data processing can usually be parallelized, with computational tasks executed simultaneously on multiple processing units to accelerate the processing process. The required programming language should also provide support for parallel computing to make full use of multi-core processors and parallel computing resources. The high performance, reliability, and rich system-level programming functions of C++ make it the preferred language, and C++ is widely used in the geophysical field. With the progress of seismic exploration technology, multi-source acquisition and high-density acquisition have gradually become common, and the scale of seismic data is growing rapidly. Seismic data processing increasingly requires the use of big data frameworks to process and analyze this massive amount of data. PySpark is the Python API of Spark, which provides an easy-to-use and understand programming interface. Compared with the traditional C++ language, Python is more concise in syntax and has higher development efficiency. Using PySpark, code can be quickly written and debugged, thus accelerating the development process of seismic data processing algorithms. By using PySpark, the distributed computing capabilities of Spark can be fully utilized to achieve parallel processing and distributed algorithms, thereby obtaining faster processing speeds on large-scale seismic data sets. As part of the Spark ecosystem, PySpark can seamlessly integrate other Spark components and tools, such as Spark SQL, Spark Streaming, and MLlib, etc. Python has rich scientific computing and data processing libraries. With the powerful ecosystems of Spark and Python, seismic data can be processed and analyzed more flexibly.

[0003] Based on this technical background, the present invention studies a method and device for batch processing of seismic data based on PySpark. Summary of the Invention

[0004] Aiming at the deficiencies of the prior art, the present invention provides a method and device for batch processing of seismic data based on PySpark. This method realizes parallel processing of the calculation process by storing segy data in blocks, defining the core function of data processing, and constructing a module development template. At the same time, by utilizing the cluster computing capabilities of Spark, the computing efficiency and processing speed are improved.

[0005] To achieve the above object, a first aspect of the present invention provides a method for batch processing seismic data based on PySpark, including:

[0006] Loading segy data using the segyio library;

[0007] Storing the segy data in chunks into different RDDs;

[0008] Defining, by a developer, a core function for data processing for each RDD block;

[0009] Constructing a module development template based on the core function for data processing for each RDD block;

[0010] Outputting and saving the processing result of the module development template.

[0011] A second aspect of the present invention provides a device for batch processing seismic data based on PySpark, including:

[0012] A loading module for loading segy data using the segyio library;

[0013] A storage module for storing the segy data in chunks into different RDDs;

[0014] A definition module for defining, by a developer, a core function of a data processing algorithm for each RDD block;

[0015] A construction module for constructing a module development template based on the core function of the data processing algorithm for each RDD block;

[0016] A saving module for outputting and saving the processing result of the module development template.

[0017] A third aspect of the present invention provides an electronic device, the electronic device including:

[0018] A memory storing executable instructions;

[0019] A processor that runs the executable instructions in the memory to implement the method for batch processing seismic data based on PySpark described in the first aspect.

[0020] A fourth aspect of the present invention provides a computer-readable storage medium that stores a computer program, and when the computer program is executed by a processor, it implements the method for batch processing seismic data based on PySpark described in the first aspect.

[0021] The beneficial effects of the present invention include:

[0022] (1) The seismic data batch processing method based on PySpark proposed by the present invention realizes parallel processing of the calculation process by storing segy data in chunks, defining the core function for data processing, and constructing a module development template. At the same time, by utilizing the cluster computing power of Spark, the calculation efficiency and processing speed are improved.

[0023] (2) The seismic data batch processing method based on PySpark proposed by the present invention is based on the seismic data batch processing framework of PySpark. By leveraging the rich scientific computing and data processing libraries of Python in seismic data processing, the input and output streams of the seismic data batch processing framework are constructed. Through the use of the cluster computing power of Spark, parallel processing of the calculation process is realized, enabling developers to focus on the implementation of data processing operations, thereby improving development efficiency and data processing efficiency.

[0024] Other features and advantages of the present invention will be described in detail in the following specific implementation section. Brief Description of the Drawings

[0025] By describing the exemplary embodiments of the present invention in more detail in conjunction with the drawings, the above and other objects, features, and advantages of the present invention will become more apparent.

[0026] Figure 1 It is a schematic flow chart of the seismic data batch processing method based on PySpark proposed by the present invention.

[0027] Figure 2 It is a schematic diagram of a directed acyclic graph of the calculation path in the Spark framework in a specific implementation manner of the seismic data batch processing method based on PySpark proposed by the present invention.

[0028] Figure 3 It is a schematic diagram of distributed computing in a specific implementation manner of the seismic data batch processing method based on PySpark proposed by the present invention.

[0029] Figure 4 It is a schematic diagram of the Spark calculation process in a specific implementation manner of the seismic data batch processing method based on PySpark proposed by the present invention. Detailed Description of the Preferred Embodiments

[0030] The following will describe the preferred embodiments of the present invention in more detail. Although the following describes the preferred embodiments of the present invention, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein.

[0031] The present invention provides a seismic data batch processing method based on PySpark, as Figure 1 shown, including:

[0032] Load SEGY data using the segyio library;

[0033] Store the SEGY data in chunks into different RDDs;

[0034] Developers define the core function for data processing on each RDD chunk;

[0035] Build a module development template based on the core function for data processing on each RDD chunk;

[0036] Output and save the processing results of the module development template.

[0037] In the present invention, through storing the SEGY data in chunks, defining the core function for data processing, and building the module development template, parallel processing of the calculation process is achieved. At the same time, by utilizing the cluster computing power of Spark, the calculation efficiency and processing speed are improved.

[0038] According to the present invention, loading SEGY data using the segyio library includes:

[0039] Load the SEGY file using the segyio library and convert it into a Spark DataFrame or RDD, or use the functions provided by the segyio library to read the SEGY data block by block and create a DataFrame or RDD containing the corresponding fields.

[0040] According to the present invention, storing the SEGY data in chunks into different RDDs includes:

[0041] Based on the chunking method of the SEGY data, call the partitioning function of Spark to store the SEGY data in chunks into different RDDs.

[0042] According to the present invention, the chunking method is to chunk according to the seismic data type by line trace number or shot number;

[0043] The partitioning function is a spatial partitioning function.

[0044] Preferably, the core function includes a seismic data processing algorithm that hopes to perform operations on the data;

[0045] The core function uses the transformation and operation functions of Spark to process the RDD chunks.

[0046] Preferably, the module development template includes a general processing function;

[0047] The general processing function divides the RDD into chunks and accepts the RDD and the core function as parameters.

[0048] According to the present invention, the general processing function divides the RDD into chunks according to seismic traces;

[0049] Output and save the processing result of the module development template, including:

[0050] Convert the processing result of the module development template into a DataFrame or RDD, create a new segy file using the segyio library, write the content in the DataFrame or RDD into the new segy file using the segyio library. After writing all the processing results, call the function flush() to flush the data into the segy file and then close the file.

[0051] In the present invention, for the seismic data batch processing framework based on PySpark, by virtue of the rich scientific computing and data processing libraries of Python in seismic data processing, the input and output streams of the seismic data batch processing framework are constructed. Through the use of the cluster computing power of Spark, the parallel processing of the computing process is realized, enabling developers to focus on the implementation of data processing operations, thereby improving the development efficiency and data processing efficiency.

[0052] The present invention will be described in more detail below through embodiments.

[0053] Embodiment 1:

[0054] In this embodiment, the seismic data batch processing framework of PySpark is first studied and analyzed, and the process is as follows:

[0055] 1) The segyio library in Python plays an important role in seismic data processing. The segyio provides functions for reading and parsing SEGY files (the standard format of seismic exploration data); it can efficiently read large SEGY files. Through memory mapping technology, it can avoid loading the entire file into memory, saving memory space and reducing the time consumption of data reading and processing; through segyio, key data such as seismic records, gathers, profiles, and seismic attributes can be conveniently accessed;

[0056] 2) The Spark framework has advantages such as high performance, scalability, multi-language support, a unified data processing framework, integration of real-time stream processing and batch processing, and a rich ecosystem and integration support in big data processing; these features make Spark the preferred framework for processing large-scale data, improving data processing efficiency, flexibility, and reliability;

[0057] 3) By means of the segyio library, the input and output of the Spark framework are transformed, enabling segy data to be stored in different RDDs in a suitable chunking manner, forming a template for the processing process of the RDD, and retaining the ability to customize the core processing function. Developers can focus on the core processing function to implement specific processing algorithms; Figure 2The figure shows a schematic diagram of a directed acyclic graph of the data calculation path in the Spark framework;

[0058] 4) Each RDD can be operated on through a series of operators. The implementation of the core processing code is to call and combine these operations and combine specific processing algorithms to achieve these purposes; the processing process will be executed on different computing nodes to achieve distributed computing, as Figure 3 shown;

[0059] 5) The overall calculation process of the Spark framework is as Figure 4 shown. After transforming the input and output of the Spark framework with the help of the segyio library and adding the core processing algorithm code, a DAG (directed acyclic graph) about the calculation path is constructed to describe when or in which processing links the calculation will occur;

[0060] 5) Spark provides a web interface (Spark WebUI) that can monitor and debug running Spark applications; by accessing the Spark WebUI through a browser, you can view the running status of jobs, the execution information of job stages, task progress, resource usage, etc.; the Spark WebUI also provides detailed information such as performance metrics and error logs to help developers analyze and optimize jobs;

[0061] Based on the above research and analysis, this embodiment provides a seismic data batch processing method based on PySpark, and the specific steps are as follows:

[0062] 1) Load segy data using the segyio library: Use the segyio library to load segy files and convert them into Spark DataFrames or RDDs; in this embodiment, the functions provided by the segyio library are used to read segy data block by block and create DataFrames or RDDs with appropriate fields;

[0063] 2) Store the segy data in different RDDs in chunks; according to your way of chunking the data, you can call the partitioning function of Spark to store the data in different RDDs in chunks; in this embodiment, according to the seismic data type, chunking is performed by line trace number or shot number, and the spatial partitioning function of Spark can be used to distribute the data chunks to different RDDs;

[0064] 3) Define the core function of the data processing algorithm by the developer; the developer can define the core processing function for each RDD block; this function should contain specific seismic data processing algorithms to complete the operations you hope to perform on the data. In this embodiment, the operations performed on the data include filtering and denoising processes; in the core processing function, Spark's transformation and operation functions can be used to process the RDD blocks;

[0065] 4) Build a module development template; to achieve a templated development process, a general processing function can be created in the development template, which accepts an RDD and the core processing function as parameters; inside this function, the RDD can be divided into blocks, and the core processing function can be called for each block for processing; in this embodiment, the RDD is divided according to seismic traces, and the core processing function (i.e., the algorithm implementation function) is called for each trace of data for processing. The developer can implement a custom core processing function according to their own needs and algorithms and pass it to the processing template function;

[0066] 5) Output and save the processing results; convert the processing results of the Spark framework into a DataFrame or RDD, use the segyio library to create a new segy file, and write the content in the DataFrame or RDD into the newly created segy file using the segyio library. After writing all the processing results, ensure that the data is flushed to the segy file by calling the flush() function and close the file.

[0067] This embodiment is based on a seismic data batch processing framework of PySpark, constructs the input and output streams of the seismic data batch processing framework with the help of rich scientific computing and data processing libraries of Python in seismic data processing, realizes parallel processing of the calculation process through the use of the cluster computing ability of Spark, enables developers to focus on the implementation of data processing operations. The Spark Web UI provides developers with a visual job monitoring function, as well as detailed information such as performance metrics and error logs, which is of great significance for improving development efficiency and data processing efficiency.

[0068] Embodiment 2:

[0069] As Figure 1 shown, a seismic data batch processing method based on PySpark in this embodiment is as follows:

[0070] Use the segyio library to load segy data;

[0071] Store the segy data in different RDDs in chunks;

[0072] Define the core function of the data processing for each RDD block by the developer;

[0073] Development template for constructing the core function of data processing based on each RDD block;

[0074] Output and save the processing results of the module development template;

[0075] In this embodiment, loading segy data using the segyio library includes:

[0076] Use the segyio library to load a segy file and convert it into a Spark DataFrame or RDD, or use the functions provided by the segyio library to read segy data block by block and create a DataFrame or RDD containing the corresponding fields;

[0077] In this embodiment, storing segy data in chunks into different RDDs includes:

[0078] Based on the chunking method of segy data, call Spark's partitioning function to store segy data in chunks into different RDDs;

[0079] In this embodiment, the chunking method is to chunk according to the seismic data type by line trace number or shot number;

[0080] The partitioning function is a spatial partitioning function;

[0081] The core function contains a seismic data processing algorithm that hopes to perform operations on the data;

[0082] The core function uses Spark's transformation and operation functions to process RDD blocks;

[0083] The module development template includes a general processing function;

[0084] The general processing function divides the RDD into chunks and accepts the RDD and the core function as parameters;

[0085] The general processing function divides the RDD into chunks according to seismic traces;

[0086] In this embodiment, outputting and saving the processing results of the module development template includes:

[0087] Convert the processing results of the module development template into a DataFrame or RDD, use the segyio library to create a new segy file, write the content in the DataFrame or RDD into the new segy file using the segyio library, and after writing all the processing results, flush the data into the segy file by calling the function flush() and then close the file.

[0088] Embodiment III:

[0089] This embodiment provides a seismic data batch processing device based on PySpark, including:

[0090] A loading module for loading segy data using the segyio library;

[0091] A storage module for storing segy data in chunks into different RDDs;

[0092] A definition module for a developer to define the core function of the data processing algorithm for each RDD block;

[0093] A construction module for constructing a module development template based on the core function of the data processing algorithm for each RDD block;

[0094] A saving module for outputting and saving the processing results of the module development template.

[0095] In this embodiment, loading segy data using the segyio library includes:

[0096] Loading a segy file using the segyio library and converting it into a Spark DataFrame or RDD, or reading segy data block by block using the functions provided by the segyio library and creating a DataFrame or RDD containing corresponding fields;

[0097] In this embodiment, storing segy data in chunks into different RDDs includes:

[0098] Based on the chunking method of segy data, calling the partitioning function of Spark to store segy data in chunks into different RDDs;

[0099] In this embodiment, the chunking method is to chunk according to the seismic data type by line trace number or shot number;

[0100] The partitioning function is a spatial partitioning function;

[0101] The core function includes a seismic data processing algorithm that hopes to perform operations on the data;

[0102] The core function uses the transformation and operation functions of Spark to process the RDD blocks;

[0103] The module development template includes a general processing function;

[0104] The general processing function divides the RDD into chunks and accepts the RDD and the core function as parameters;

[0105] The general processing function divides the RDD into chunks according to seismic traces;

[0106] In this embodiment, the output and saving of the processing result of the module development template include:

[0107] Convert the processing result of the module development template into a DataFrame or RDD, use the segyio library to create a new segy file, write the content in the DataFrame or RDD into the new segy file using the segyio library, and after writing all the processing results, flush the data into the segy file by calling the function flush() and then close the file.

[0108] Embodiment 4:

[0109] An embodiment of the present invention provides an electronic device including a memory and a processor.

[0110] The memory stores executable instructions.

[0111] The processor runs the executable instructions in the memory to implement a seismic data batch processing method based on PySpark.

[0112] The memory is used to store non-temporary computer-readable instructions. Specifically, the memory may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.

[0113] The processor may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions. In an embodiment of the present invention, the processor is used to run the computer-readable instructions stored in the memory.

[0114] Those skilled in the art should understand that in order to solve the technical problem of how to obtain a good user experience effect, this embodiment may also include well-known structures such as communication buses and interfaces, and these well-known structures should also be included in the protection scope of the present invention.

[0115] For the detailed description of this embodiment, reference may be made to the corresponding descriptions in the foregoing embodiments, and details will not be repeated here.

[0116] Embodiment 5:

[0117] An embodiment of the present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements a seismic data batch processing method based on PySpark.

[0118] According to the computer-readable storage medium of the embodiments of the present invention, non-temporary computer-readable instructions are stored thereon. When the non-temporary computer-readable instructions are run by a processor, all or part of the steps of the methods of the various embodiments of the present invention described above are executed.

[0119] The above computer-readable storage medium includes but is not limited to: optical storage media (such as CD-ROMs and DVDs), magneto-optical storage media (such as MOs), magnetic storage media (such as magnetic tapes or external hard drives), media with built-in rewritable non-volatile memories (such as memory cards), and media with built-in ROMs (such as ROM cartridges).

[0120] The seismic data batch processing method based on PySpark proposed by the embodiments of the present invention realizes parallel processing of the calculation process by storing segy data in blocks, defining the core functions of data processing, and constructing a module development template. At the same time, by utilizing the cluster computing power of Spark, the computing efficiency and processing speed are improved.

[0121] The various embodiments of the present invention have been described above. The above description is exemplary and not exhaustive, and is also not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments.

Claims

1. A method for batch processing of seismic data based on PySpark, characterized in that, it includes: Loading segy data using the segyio library; Storing the segy data in different RDDs in chunks; Defining by the developer the core function for data processing for each RDD block; Constructing a module development template based on the core function for data processing for each RDD block; Outputting and saving the processing results of the module development template.

2. The method according to claim 1, characterized in that, Loading segy data using the segyio library includes: Loading a segy file using the segyio library and converting it into a Spark DataFrame or RDD, or reading the segy data block by block using the functions provided by the segyio library and creating a DataFrame or RDD containing the corresponding fields.

3. The method according to claim 2, characterized in that, Storing the segy data in different RDDs in chunks includes: Based on the chunking method of the segy data, calling the partitioning function of Spark to store the segy data in different RDDs in chunks.

4. The method according to claim 3, characterized in that, The chunking method is to chunk according to the seismic data type by line trace number or shot number; The partitioning function is a spatial partitioning function.

5. The method according to claim 3, characterized in that, The core function contains a seismic data processing algorithm that hopes to perform operations on the data; The core function uses the transformation and operation functions of Spark to process the RDD blocks.

6. The method according to claim 5, characterized in that, The module development template includes a general processing function; The general processing function divides the RDD into blocks and accepts the RDD and the core function as parameters.

7. The method according to claim 6, characterized in that, The general processing function divides the RDD into blocks according to seismic traces; Outputting and saving the processing results of the module development template includes: Converting the processing results of the module development template into a DataFrame or RDD, using the segyio library to create a new segy file, writing the content in the DataFrame or RDD into the new segy file using the segyio library, and after writing all the processing results, flushing the data into the segy file by calling the function flush() and then closing the file.

8. An apparatus for batch processing of seismic data based on PySpark, characterized in that, it includes: A loading module for loading segy data using the segyio library; A storage module for storing the segy data in different RDDs in chunks; A definition module for the developer to define the core function of the data processing algorithm for each RDD block; A construction module for constructing a module development template based on the core function of the data processing algorithm for each RDD block; A saving module for outputting and saving the processing results of the module development template.

9. An electronic device, characterized in that, The electronic device includes: a memory storing executable instructions; a processor that runs the executable instructions in the memory to implement the PySpark-based seismic data batch processing method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, the computer-readable storage medium stores a computer program which, when executed by a processor, implements the PySpark-based seismic data batch processing method according to any one of claims 1-7.

Citation Information

Cited By

  • Script language-oriented seismic batch processing job editing method and device

    CN120892030A

  • Script language oriented earthquake batch job editing method and device

    CN120892030B