Method for quickly sequencing seismic data

By segmenting seismic data files and extracting and sorting information with key-value, the problem of low efficiency in the process of seismic data is solved, and the rapid sorting of massive seismic data is achieved, which significantly saves space and time and improves the analysis and processing efficiency.

CN120122169APending Publication Date: 2025-06-10CHINA PETROLEUM & CHEMICAL CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311685333.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-08
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

In the processing of seismic data, the original data is not processed in an orderly manner, which leads to huge challenges in the later track set extraction and sorting work, especially for massive earthquake data (hundreds of G or even T levels), the processing efficiency is low and the space and time occupies a lot.

Method used

By determining the segmentation particle size of the earthquake data file, the data file is divided into multiple data blocks, the key-value pair information is extracted, including the header fields and address information of the earthquake path to which the earthquake data belongs, merge it into a key-value pair information set, and sort it according to business needs, to achieve quick sorting of earthquake data.

Benefits of technology

It realizes the rapid sorting and processing of super-large seismic data files, greatly saving space and time, and improving the work efficiency of data analysis and processing. Especially on large supercomputing platforms, the efficiency is significantly improved for applications with strict storage space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120122169A_ABST
    Figure CN120122169A_ABST
Patent Text Reader

Abstract

The invention provides a method for quickly sorting seismic data. The method comprises the steps that the segmentation granularity of a seismic data file to be processed is determined, the seismic data file is divided into a plurality of data blocks based on the segmentation granularity, and seismic data, belonging to the same seismic channel, in the seismic data file are located in the same data block after division; key value pair information is extracted from the seismic data in the multiple data blocks, and the key value pair information comprises a trace header field and address information of a seismic trace to which the seismic data belongs; combining the extracted key value pair information into a key value pair information set; and sorting the key value pair information in the key value pair information set according to service requirements, and performing service processing on the seismic data in the seismic data file according to the sorted key value pair information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of seismic exploration data processing, and particularly relates to a method, device, computer-readable storage medium and electronic device for quickly sorting seismic data. Background Art

[0002] During the processing of seismic data, the required trace gather needs to be extracted according to a certain dimension, and then the ordered trace gather is analyzed and processed to form a visual output. Since the original seismic data is not sorted during the collection process, and generally seismic data is massive data, ranging from hundreds of gigabytes at least to dozens or even hundreds of terabytes, this poses a huge challenge to the later trace gather extraction and sorting work. Summary of the Invention

[0003] In view of the above problems, the present invention provides a method for quickly sorting seismic data.

[0004] In a first aspect, an embodiment of the present invention provides a method for quickly sorting seismic data, the method comprising:

[0005] Determine the segmentation granularity of the seismic data file to be processed, and divide the seismic data file into multiple data blocks based on the segmentation granularity. After division, the seismic data belonging to the same seismic trace in the seismic data file is located in the same data block;

[0006] Extract key-value pair information from the seismic data in the multiple data blocks, where the key-value pair information includes the header field and address information of the seismic trace to which the seismic data belongs;

[0007] Merge the extracted key-value pair information into a key-value pair information set;

[0008] Sort the key-value pair information in the key-value pair information set according to service requirements, and perform service processing on the seismic data in the seismic data file according to the sorted key-value pair information.

[0009] According to an embodiment of the present invention, the segmentation granularity of the seismic data file to be processed is determined according to the following formula:

[0010] E = MAX(A, C / D, B)

[0011] where E is the segmentation granularity, A is the size of the storage space for storing seismic data, B is the size of the seismic trace, C is the size of the seismic data file, and D is the parallelism for processing seismic data.

[0012] According to an embodiment of the present invention, the key-value pair information is extracted from the seismic data in the multiple data blocks in the following manner:

[0013] Extract key-value pair information in a manner that adopts parallel processing among the multiple data blocks and sequential processing within each data block.

[0014] According to an embodiment of the present invention, the address information includes a trace header offset and a trace data offset.

[0015] According to an embodiment of the present invention, the set of key-value pair information is a resilient distributed dataset.

[0016] According to an embodiment of the present invention, sorting the key-value pair information in the set of key-value pair information according to service requirements includes:

[0017] Sort the key-value pair information in the set of key-value pair information according to the trace header field.

[0018] According to an embodiment of the present invention, performing service processing on the seismic data in the seismic data file according to the sorted key-value pair information includes:

[0019] Extract and / or sort the seismic data of the target trace gather in the seismic data file according to the address information in the sorted key-value pair information.

[0020] In a second aspect, an embodiment of the present invention provides an apparatus for quickly sorting seismic data, which is characterized by including:

[0021] A splitting module, configured to determine the splitting granularity of the seismic data file to be processed, and divide the seismic data file into multiple data blocks based on the splitting granularity. After division, the seismic data belonging to the same seismic trace in the seismic data file is located in the same data block;

[0022] An extraction module, configured to extract key-value pair information from the seismic data in the multiple data blocks, where the key-value pair information includes the trace header field and address information of the seismic trace to which the seismic data belongs;

[0023] A merging module, configured to merge the extracted key-value pair information into a set of key-value pair information;

[0024] A sorting module, configured to sort the key-value pair information in the set of key-value pair information according to service requirements, and perform service processing on the seismic data in the seismic data file according to the sorted key-value pair information.

[0025] In a third aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements a method for quickly sorting seismic data as described in the first aspect above.

[0026] Fourthly, an embodiment of the present invention provides an electronic device, which includes:

[0027] a processor;

[0028] a memory for storing executable instructions of the processor;

[0029] wherein, the processor is configured to execute the instructions to implement a method for quickly sorting seismic data as described in the first aspect above.

[0030] Compared with the prior art, the above technical solution of the present invention has the following beneficial effects:

[0031] The present invention designs a method for randomly reading an original seismic data file with a variable step size, obtaining required gather information and performing quick sorting. This method can quickly perform data sorting processing on ultra-large seismic data files, especially can complete gather extraction and sorting processing of seismic data files of hundreds of gigabytes or even several terabytes in a very short time, greatly saving the space and time of seismic data processing and improving the working efficiency of subsequent data analysis and processing. Using the method provided by the present invention can greatly improve the gather extraction efficiency of seismic data, increase the speed, reduce the memory and storage occupancy, especially for applications deployed on large supercomputer platforms and with relatively strict requirements for local storage space, the efficiency improvement is obvious. Description of the Drawings

[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0033] Figure 1 is a step flow chart of the method for quickly sorting seismic data provided by the embodiment of the present invention;

[0034] Figure 2 is Figure 1 a schematic diagram of file data obtained by the method for quickly sorting seismic data;

[0035] Figure 3 is a schematic diagram of the composition structure of the electronic device provided by the embodiment of the present invention. Detailed Embodiments

[0036] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0037] Embodiment 1

[0038] Figure 1 is a flowchart of the method for quickly sorting seismic data in the embodiments of the present invention. As Figure 1 shown, the method mainly includes the following steps.

[0039] (1) Divide the seismic data file to be processed into multiple data blocks.

[0040] In this embodiment, preferably, the Spark parallel processing technology is used to process the original seismic data file.

[0041] In order to maximize the parallel processing efficiency of Spark during parallel processing, first determine the segmentation granularity of the seismic data file data. In this embodiment, the seismic data is stored on HDFS (Hadoop Distributed FileSystem), and HDFS has its own block size of A; the smallest unit of the seismic data is a trace, and the trace size is B; the total seismic data size of the seismic data file is C; the desired parallelism of the user is D, then the reasonable segmentation granularity E is calculated according to the following formula:

[0042] E = MAX(A, C / D, B)

[0043] This can ensure that after division, the seismic data belonging to the same seismic trace in the seismic data file will not be split into two different data blocks, that is, the seismic data belonging to the same seismic trace is located in the same data block, and at the same time, the maximum parallelism can be ensured.

[0044] (2) Extract key-value pair information {Key, Value} for each data block.

[0045] In this embodiment, it is preferable to extract key-value pair information in a manner of parallel processing among multiple data blocks and sequential processing within each data block. When processing within each block, the header fields of the seismic traces to which the seismic data belongs are extracted: Key(k1, k2,...), and at the same time, the header offset and data offset Value(offsetHDR, offsetDATA) that record the storage address of the seismic trace are also extracted, constituting the key-value pair information, that is: Key(k1, k2,...), Value(offsetHDR, offsetDATA).

[0046] (3) Merge the extracted key-value pair information into a set of key-value pair information.

[0047] Merge the extracted key-value pair information into a set of key-value pair information. In this embodiment, the set of key-value pair information can be the Resilient Distributed Datasets (RDD) of Spark.

[0048] (4) Sort the key-value pair information in the set of key-value pair information according to business requirements, and perform business processing on the seismic data in the seismic data file according to the sorted key-value pair information.

[0049] Specifically, write a sorting rule according to the requirements of the seismic data processing business, and perform sorting processing on the key-value pair information in the RDD of the set of key-value pair information to obtain the sorted RDD. As Figure 2 shown, in this embodiment, sort the key-value pair information according to the header fields of the seismic traces, and then perform business processing on the seismic data in the seismic data file according to the sorted key-value pair information. For example. The seismic data of the target trace set in the seismic data file can be extracted and / or sorted according to the address information in the sorted key-value pair information. Of course, in practical applications, the seismic data business processing can be various and is not limited to this.

[0050] Using the method provided in this embodiment can greatly improve the extraction efficiency of seismic data trace sets, increase the speed, reduce the memory and storage occupancy, especially for applications deployed on large supercomputer platforms with relatively strict requirements for local storage space, the efficiency improvement is obvious.

[0051] Embodiment Two

[0052] This embodiment also provides a device for quickly sorting seismic data, which includes:

[0053] A splitting module, configured to determine the splitting granularity of a seismic data file to be processed, and divide the seismic data file into multiple data blocks based on the splitting granularity. After division, the seismic data belonging to the same seismic trace in the seismic data file is located in the same data block;

[0054] An extraction module, configured to extract key-value pair information from the seismic data in the multiple data blocks, where the key-value pair information includes the header field and address information of the seismic trace to which the seismic data belongs;

[0055] A merging module, configured to merge the extracted key-value pair information into a set of key-value pair information;

[0056] A sorting module, configured to sort the key-value pair information in the set of key-value pair information according to service requirements, and perform service processing on the seismic data in the seismic data file according to the sorted key-value pair information.

[0057] Embodiment III

[0058] This embodiment also provides a computer-readable medium, on which a computer program is stored. When the program is executed by a processor, it implements each step of a method for reducing the risk of casing deformation in horizontal wells as described in the above embodiment.

[0059] It should be noted that all or part of the processes of implementing the method in the above embodiment of the present invention can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. Of course, there are other ways of readable storage media, such as quantum memory, graphene memory, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0060] Embodiment IV

[0061] Figure 2This is a schematic structural diagram of an electronic device according to an embodiment of the present invention. As Figure 2 shown, at the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and a memory. Among them, the memory may include a memory, such as a high-speed random access memory (Random-Access Memory, RAM), and may also include a non-volatile memory, such as at least one disk memory, etc. Of course, the electronic device may also include other hardware required for other services.

[0062] The processor, network interface, and memory can be interconnected through an internal bus, and the internal bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, only line segments are used in the figure, but it does not mean that there is only one bus or one type of bus.

[0063] The memory is used to store programs. Specifically, the program may include program code, and the program code includes computer operation instructions. The memory may include a memory and a non-volatile memory, and provide instructions and data to the processor. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it. The processor executes the program stored in the memory to execute all the steps in the foregoing method for reducing the risk of casing deformation in horizontal wells.

[0064] The communication bus mentioned in the above device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus. The communication interface is used for communication between the above electronic device and other devices.

[0065] The bus includes hardware, software, or both, for coupling the above components to each other. By way of example, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable bus or a combination of two or more of these. Where appropriate, the bus may include one or more buses. Although embodiments of the present invention describe and illustrate specific buses, the present invention contemplates any suitable bus or interconnect.

[0066] The memory may include Random Access Memory (RAM), or may also include non-volatile memory (Non-Volatile Memory, NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.

[0067] The memory may include mass storage for data or instructions. By way of example and not limitation, the memory may include a Hard Disk Drive (HDD), a floppy disk drive, flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory may include removable or non-removable (or fixed) media. In a particular embodiment, the memory is non-volatile solid state memory. In a particular embodiment, the memory includes Read Only Memory (ROM). Where appropriate, the ROM may be a mask-programmed ROM, a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically Erasable PROM (EEPROM), an Electrically Alterable ROM (EAROM), or flash memory, or a combination of two or more of these.

[0068] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0069] It should be noted that those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present invention. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiment and will not be repeated here.

[0070] The device, equipment, system, module or unit illustrated in the above embodiments can be specifically implemented by a computer chip or an entity, or by a product with a certain function. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, an in-vehicle human-machine interaction device, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0071] Although the present invention provides method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-creative means. The order of steps listed in the embodiments is only one way among the execution orders of numerous steps and does not represent the only execution order. When the actual device or terminal product executes, it can be executed in the order of the method shown in the embodiments or the drawings or executed in parallel (for example, in a parallel processor or multi-threaded processing environment, or even in a distributed data processing environment).

[0072] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing device to produce a machine such that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.

[0073] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.

[0074] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.

[0075] It should be noted that, in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0076] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the device, electronic device, and readable storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the corresponding parts of the method embodiments for the relevant content.

[0077] The above are only the preferred embodiments of the present invention, and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are all included in the protection scope of the present invention.

Claims

1. A method for quickly sorting seismic data, characterized in that, it includes the following steps: S100, determining the segmentation granularity of the seismic data file to be processed, and dividing the seismic data file into multiple data blocks based on the segmentation granularity. After division, the seismic data belonging to the same seismic trace in the seismic data file is located in the same data block; S200, extracting key-value pair information from the seismic data in the multiple data blocks, where the key-value pair information includes the header field and address information of the seismic trace to which the seismic data belongs; S300, merging the extracted key-value pair information into a key-value pair information set; S400, sorting the key-value pair information in the key-value pair information set according to business requirements, and performing business processing on the seismic data in the seismic data file according to the sorted key-value pair information.

2. The method according to claim 1, characterized in that, the segmentation granularity of the seismic data file to be processed is determined according to the following formula: E = MAX(A, C / D, B) where E is the segmentation granularity, A is the size of the storage space for storing seismic data, B is the size of the seismic trace, C is the size of the seismic data file, and D is the parallelism for processing seismic data.

3. The method according to claim 1, characterized in that, the key-value pair information is extracted from the seismic data in the multiple data blocks in the following manner: The key-value pair information is extracted in a manner of parallel processing among the multiple data blocks and sequential processing within each data block.

4. The method according to claim 1, characterized in that, the address information includes the header offset and the trace data offset.

5. The method according to claim 1, characterized in that, the key-value pair information set is an elastic distributed dataset.

6. The method according to claim 1, characterized in that, the sorting of the key-value pair information in the key-value pair information set according to business requirements includes: sorting the key-value pair information in the key-value pair information set according to the header field.

7. The method according to claim 1, characterized in that, the business processing of the seismic data in the seismic data file according to the sorted key-value pair information includes: extracting and / or sorting the seismic data of the target trace set in the seismic data file according to the address information in the sorted key-value pair information.

8. An apparatus for quickly sorting seismic data, characterized in that, it includes: a segmentation module, configured to determine the segmentation granularity of the seismic data file to be processed, and divide the seismic data file into multiple data blocks based on the segmentation granularity. After division, the seismic data belonging to the same seismic trace in the seismic data file is located in the same data block; an extraction module, configured to extract key-value pair information from the seismic data in the multiple data blocks, where the key-value pair information includes the header field and address information of the seismic trace to which the seismic data belongs; a merging module, configured to merge the extracted key-value pair information into a key-value pair information set; A sorting module, configured to sort the key-value pair information in the key-value pair information set according to service requirements, and perform service processing on the seismic data in the seismic data file according to the sorted key-value pair information.

9. A computer-readable storage medium, characterized in that, a computer program is stored thereon, and when the program is executed by a processor, it implements a method for quickly sorting seismic data as described in any one of claims 1 to 7.

10. An electronic device, comprising: a processor; a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the instructions to implement a method for quickly sorting seismic data as described in any one of claims 1 to 7.