Seismic data processing method and device based on Spark framework
Through the seismic data processing method based on the Spark framework, the problem of slow processing speed caused by data overlap in seismic data processing is solved, and efficient data processing and labor costs are achieved.
Patent Information
- Application Number
- CN202311686726.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-08
- Publication Date
- 2025-06-10
AI Technical Summary
In seismic data processing, because the data in the window area overlaps with adjacent window data, the overall data volume of seismic data is large and the overall processing speed is slow.
Using the seismic data processing method based on the Spark framework, multiple window-range data sets are obtained by dividing and processing the seismic data, and a first elastic distributed data set is constructed based on the input range and output range of each window-range data set, thereby generating the second elastic distributed data set. Finally, the second elastic distributed data set is integrated and sorted in the preset order.
Through parallel processing and automation window range setting, the time and labor cost of seismic data processing are significantly reduced and processing efficiency is improved.
Smart Images

Figure CN120122174A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of seismic data processing in the field of earth sciences, and more specifically, to a method and apparatus for processing seismic data based on the Spark framework. Background Art
[0002] In seismic exploration, it is sometimes necessary to perform window scanning processing on seismic data. There is overlap between the data within the window area and the data of adjacent windows. This way of data overlap results in a large overall data volume of the seismic data and a slow overall processing speed. Therefore, there is an urgent need for a method that can improve the processing efficiency of seismic data to solve the above problems. Summary of the Invention
[0003] In view of this, the present invention discloses a technical solution for improving the processing speed of seismic data. According to one aspect of the present invention, a method for processing seismic data based on the Spark framework is proposed. The method is applied to the Spark framework, and the processing method includes: dividing the seismic data to obtain multiple window range data sets; constructing a first resilient distributed data set according to the input range and the corresponding output range corresponding to each window range data set; generating a second resilient distributed data set according to the first resilient distributed data set; wherein, the second resilient distributed data set is a key-value pair data set obtained by computational reconstruction from the first resilient distributed data set, the key corresponds to the line number and the trace number, and the value corresponds to the single-trace data; sorting and integrating the data of the second resilient distributed data set in accordance with a preset order between multiple line numbers and a preset order between multiple trace numbers corresponding to each line number, and using it as the processing result.
[0004] In some embodiments, the dividing the seismic data to obtain multiple window range data sets includes: obtaining the input parameter size and the output parameter size; sequentially sampling the seismic data according to the input parameter size to obtain multiple window range data sets, the input range corresponding to each window range data set, and the output range corresponding to each window range data set.
[0005] In some embodiments, the generating a second resilient distributed data set according to the first resilient distributed data set includes: obtaining the input data corresponding to each input range in parallel according to the multiple input ranges in the first resilient distributed data set; generating an array formed by multiple single-trace data corresponding to each input data in parallel according to the input data corresponding to each input range and the output range corresponding to each input range; separating each single-trace data according to the byte array size of the stored single-trace data, and reconstructing it into a key-value pair form as the second resilient distributed data set; wherein, the key is a structure composed of the line number and the trace number, and the value is an array of single-trace data bodies.
[0006] In some embodiments, integrating and sorting the second resilient distributed dataset according to a preset order among multiple line numbers and a preset order among multiple trace numbers corresponding to each line number, and using the result as the processing result, includes: sequentially obtaining single-trace data of each trace number corresponding to each line number in the second resilient distributed dataset according to the preset order among multiple trace numbers corresponding to each line number, and integrating the sequentially obtained multiple single-trace data into a trace gather corresponding to the line number; sorting the trace gathers corresponding to each line number according to the preset order among multiple line numbers, and using the result as the processing result.
[0007] In some embodiments, the input range includes an input start line number, an input end line number, an input start trace number, and an input end trace number; the output range includes: an output start line number, an output end line number, an output start trace number, and an output end trace number.
[0008] According to one aspect of the present invention, there is also provided a processing device for seismic data based on the Spark framework, which is applied to the Spark framework. The processing device includes: a data partitioning module for partitioning seismic data to obtain multiple window range datasets; a first resilient distributed dataset construction module for constructing a first resilient distributed dataset according to the input range and the corresponding output range corresponding to each window range dataset; a second resilient distributed dataset construction module for generating a second resilient distributed dataset according to the first resilient distributed dataset; wherein, the second resilient distributed dataset is a key-value pair dataset obtained by computational reconstruction from the first resilient distributed dataset, the key corresponds to the line number and the trace number, and the value corresponds to the single-trace data; a processing result generation module for integrating and sorting the second resilient distributed dataset according to a preset order among multiple line numbers and a preset order among multiple trace numbers corresponding to each line number, and using the result as the processing result.
[0009] In some embodiments, the data partitioning module is further configured to obtain the input parameter size and the output parameter size; sequentially sample the seismic data according to the input parameter size to obtain multiple window range datasets, the input range corresponding to each window range dataset, and the output range corresponding to each window range dataset.
[0010] In some embodiments, the second resilient distributed dataset construction module is further configured to obtain the input data corresponding to each input range in parallel according to a plurality of input ranges in the first resilient distributed dataset; generate, in parallel, an array formed by a plurality of single-trace data corresponding to each input data according to the input data corresponding to each input range and the output range corresponding to each input range; separate each single-trace data according to the byte array size of the single-trace data storage, and reconstruct it into a key-value pair form as the second resilient distributed dataset; wherein the key is a structure composed of a line number and a trace number, and the value is an array of single-trace data bodies.
[0011] In some embodiments, the processing result generation module is further configured to sequentially obtain the single-trace data of each trace number corresponding to each line number in the second resilient distributed dataset according to a preset order among the plurality of trace numbers corresponding to each line number, and integrate the sequentially obtained plurality of single-trace data into a trace gather corresponding to the line number; sort the trace gathers corresponding to each line number according to a preset order among the plurality of line numbers, and use it as the processing result.
[0012] In some embodiments, the input range includes an input start line number, an input end line number, an input start trace number, and an input end trace number; the output range includes: an output start line number, an output end line number, an output start trace number, and an output end trace number.
[0013] According to another aspect of the present invention, an electronic device is further provided. The electronic device includes: a memory storing executable instructions; a processor that runs the executable instructions in the memory to implement the above-mentioned method for processing seismic data based on the Spark framework.
[0014] According to another aspect of the present invention, a computer-readable storage medium is further provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned method for processing seismic data based on the Spark framework.
[0015] The technical solution has at least the following advantages: A method for processing seismic data based on the Spark framework provided by an embodiment of the present invention can perform partitioning processing on seismic data to obtain multiple window range data sets, and then construct a first resilient distributed data set according to the corresponding input range and output range of each window range data set. Then, based on the first resilient distributed data set, a second resilient distributed data set is generated. Finally, according to the preset order between multiple line numbers and the preset order between multiple trace numbers corresponding to each line number, data integration and sorting are performed on the second resilient distributed data set, which is used as the processing result. The processing method provided by the embodiment of the present invention defines the window range in the first resilient distributed data set by setting the input range and output range, so as to replace manual partitioning, which can reduce labor costs and manual time consumption. When developers use this set of frameworks, they only need to implement the core algorithm part, and the rest can be completed by the framework. In addition, in the process of generating the second resilient distributed data set from the first resilient distributed data set, parallel processing can be achieved relying on the Spark framework, thereby greatly reducing the processing time.
[0016] The methods and apparatuses of the present invention have other characteristics and advantages, which will be obvious in the accompanying drawings and subsequent detailed implementation manners incorporated herein, or will be described in detail in the accompanying drawings and subsequent detailed implementation manners incorporated herein. These accompanying drawings and detailed implementation manners are used together to explain the specific principles of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] By describing the exemplary embodiments of the present invention in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present invention will become more obvious. Among them, in the exemplary embodiments of the present invention, the same reference numerals generally represent the same components.
[0018] Figure 1 The flowchart of a method for processing seismic data based on the Spark framework according to an embodiment of the present invention is shown.
[0019] Figure 2 The reference schematic diagram of a method for processing seismic data based on the Spark framework according to an embodiment of the present invention is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] In the prior art, the general method for processing seismic data is to calculate the overall input trace gather (a trace gather includes multiple single-trace data, and single-trace data or single-trace seismic data can be referred to in related technologies), and then output the result trace gather. The whole process is a serial process. In addition, due to the large amount of seismic data, the whole processing process usually takes a lot of time. Moreover, for the window range data set with data overlapping parts obtained by window scanning processing, developers usually need to manually set the overlapping range, and the labor cost and processing time are also relatively high.
[0021] In view of this, the embodiments of the present invention provide a method for processing seismic data based on the Spark framework. The seismic data can be divided and processed to obtain multiple window range data sets. Then, according to the corresponding input range and output range of each window range data set, a first resilient distributed data set is constructed. Then, based on the first resilient distributed data set, a second resilient distributed data set is generated. Finally, according to the preset order among multiple line numbers and the preset order among multiple trace numbers corresponding to each line number, the data in the second resilient distributed data set is integrated and sorted, which is used as the processing result. The processing method provided by the embodiments of the present invention defines the window range in the first resilient distributed data set by setting the input range and output range, so as to replace manual division, which can reduce labor costs and manual time consumption. In addition, in the process of generating the second resilient distributed data set from the first resilient distributed data set, parallel processing can be achieved relying on the Spark framework, thereby greatly reducing the processing time.
[0022] Specifically, the present invention proposes a method for processing seismic data based on the Spark framework, which is applied to the Spark framework. The processing method includes: dividing and processing the seismic data to obtain multiple window range data sets; constructing a first resilient distributed data set according to the corresponding input range and output range of each window range data set; generating a second resilient distributed data set according to the first resilient distributed data set; wherein, the second resilient distributed data set is a key-value pair data set obtained by calculating and reconstructing the first resilient distributed data set, the key corresponds to the line number and trace number, and the value corresponds to the single-trace data; integrating and sorting the data in the second resilient distributed data set according to the preset order among multiple line numbers and the preset order among multiple trace numbers corresponding to each line number, and using it as the processing result.
[0023] In some embodiments, the dividing and processing the seismic data to obtain multiple window range data sets includes: obtaining the input parameter size and the output parameter size; sampling the seismic data sequentially according to the input parameter size to obtain multiple window range data sets, the corresponding input range of each window range data set, and the corresponding output range of each window range data set.
[0024] In some embodiments, generating a second resilient distributed dataset according to the first resilient distributed dataset includes: obtaining input data corresponding to each input range in parallel according to a plurality of input ranges in the first resilient distributed dataset; generating, in parallel according to the input data corresponding to each input range and the output range corresponding to each input range, an array formed by a plurality of single-trace data corresponding to each input data; separating each single-trace data according to the size of the byte array stored in the single-trace data, and reconstructing it into a key-value pair form as the second resilient distributed dataset; wherein the key is a structure composed of a line number and a trace number, and the value is an array of single-trace data bodies.
[0025] In some embodiments, performing data integration and sorting on the second resilient distributed dataset according to a preset order between a plurality of line numbers and a preset order between a plurality of trace numbers corresponding to each line number, and using it as a processing result includes: sequentially obtaining the single-trace data of each trace number corresponding to each line number in the second resilient distributed dataset according to the preset order between the plurality of trace numbers corresponding to each line number, and integrating the sequentially obtained plurality of single-trace data into a trace gather corresponding to the line number; sorting the trace gathers corresponding to each line number according to the preset order between the plurality of line numbers, and using it as the processing result.
[0026] In some embodiments, the input range includes an input start line number, an input end line number, an input start trace number, and an input end trace number; the output range includes: an output start line number, an output end line number, an output start trace number, and an output end trace number.
[0027] The present invention also provides a processing device for seismic data based on the Spark framework, which is applied to the Spark framework. The processing device includes: a data partitioning module for partitioning seismic data to obtain a plurality of window range datasets; a first resilient distributed dataset construction module for constructing a first resilient distributed dataset according to the input range and the corresponding output range corresponding to each window range dataset; a second resilient distributed dataset construction module for generating a second resilient distributed dataset according to the first resilient distributed dataset; wherein the second resilient distributed dataset is a key-value pair dataset obtained by computational reconstruction from the first resilient distributed dataset, the key corresponds to the line number and the trace number, and the value corresponds to the single-trace data; a processing result generation module for performing data integration and sorting on the second resilient distributed dataset according to a preset order between a plurality of line numbers and a preset order between a plurality of trace numbers corresponding to each line number, and using it as the processing result.
[0028] In some embodiments, the data partitioning module is further configured to obtain the sizes of the input parameters and the output parameters; and sample the seismic data in sequence according to the size of the input parameters to obtain a plurality of window range data sets, the input range corresponding to each window range data set, and the output range corresponding to each window range data set.
[0029] In some embodiments, the second resilient distributed dataset construction module is further configured to obtain, in parallel, the input data corresponding to each input range according to the multiple input ranges in the first resilient distributed dataset; generate, in parallel, an array formed by a plurality of single-trace data corresponding to each input data according to the input data corresponding to each input range and the output range corresponding to each input range; separate each single-trace data according to the size of the byte array stored in the single-trace data, and reconstruct it into a key-value pair form as the second resilient distributed dataset; wherein the key is a structure composed of a line number and a trace number, and the value is an array of single-trace data bodies.
[0030] In some embodiments, the processing result generation module is further configured to sequentially obtain the single-trace data of each trace number corresponding to each line number in the second resilient distributed dataset according to the preset order among the multiple trace numbers corresponding to each line number, and integrate the sequentially obtained multiple single-trace data into a trace gather corresponding to the line number; sort the trace gathers corresponding to each line number according to the preset order among the multiple line numbers, and use it as the processing result.
[0031] In some embodiments, the input range includes the starting line number of the input, the ending line number of the input, the starting trace number of the input, and the ending trace number of the input; the output range includes: the starting line number of the output, the ending line number of the output, the starting trace number of the output, and the ending trace number of the output.
[0032] The present invention also provides an electronic device, which includes: a memory storing executable instructions; a processor, where the processor runs the executable instructions in the memory to implement the above-mentioned method for processing seismic data based on the Spark framework.
[0033] The present invention also provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned method for processing seismic data based on the Spark framework.
[0034] The preferred embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the preferred embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to make the present invention more thorough and complete, and to fully convey the scope of the present invention to those skilled in the art.
[0035] Example 1
[0036] Figure 1 The flowchart of a method for processing seismic data based on the Spark framework according to an embodiment of the present invention is shown. As shown in the figure, the method includes steps 1 to 4. The above-mentioned Spark framework is an open general-purpose in-memory parallel computing framework, which has a fast running speed, good ease of use, and strong versatility. For specific framework details, reference can be made to related technologies.
[0037] Step 1: Perform partitioning processing on the seismic data to obtain multiple window range data sets. Combining Figure 2 as shown, Figure 2 The reference schematic diagram of a method for processing seismic data based on the Spark framework according to an embodiment of the present invention is shown. Figure 2 In it, 1542 to 1554 are used to represent the trace numbers, and 1026 to 1036 are used to represent the line numbers. The two constitute the distribution of the seismic data. In this figure, three rectangles respectively represent three rectangular windows scanning the seismic data, and what is obtained after the scanning ends (that is, partitioning) is the window range data set. It should be understood that for the scanning process, the size of the rectangular window and the scanning rules (for example: the single sliding distance and sliding direction of the rectangular window) can be set by developers according to the actual situation, and the embodiments of the present invention do not limit this here. The window range data set can include multiple single-trace data. For the data form and composition of single-trace data and seismic data, reference can be made to related technologies. There may be data overlap between rectangular windows, that is, partial line numbers, partial trace numbers, and partial single-trace data overlap. It should be understood that if developers have specific requirements, the partitioning processing method can be changed, and the embodiments of the present invention do not limit this here.
[0038] In a possible implementation manner, step 1 may include: obtaining the input parameter size and the output parameter size. The input parameter size can be used to define the size of the above-mentioned rectangular window, and the output parameter size can be used to define the size obtained by different types of operations input by developers for the single-trace data in the rectangular window (if analogized by the rectangular window during the above scanning, it can also be presented as a rectangle in arrangement). Then, according to the input parameter size, sample the seismic data in sequence to obtain multiple window range data sets, the input range corresponding to each window range data set, and the output range corresponding to each window range data set. Exemplarily, sampling is the above-mentioned scanning. The total number of trace numbers and line numbers in the input range is the same as the input parameter size, and the total number of trace numbers and line numbers in the output range is the same as the output parameter size. The above input range and output range can both include specific trace numbers and line numbers.
[0039] Continue to refer to Figure 1, Step 2, construct a first Resilient Distributed Dataset (RDD) according to the input range and the corresponding output range of each window range dataset. Exemplarily, the above input range can be represented as a range of track numbers and a range of line numbers. For example, in Figure 2 Figure 2 the range from track number 1542 to 1546 (i.e., the above range of track numbers) and the range from line number 1033 to 1036 (i.e., the above range of line numbers) are the input ranges corresponding to a window range dataset. Similarly, the output range can be represented as a range of track numbers and a range of line numbers. In one example, the range of track numbers and the range of line numbers in the input range may be larger than those in the output range, which specifically depends on the processing method and purpose of the seismic data. This embodiment of the present invention does not limit this, that is, the numerical sizes of the range of track numbers and the range of line numbers in both the input range and the output range can be different. The Resilient Distributed Dataset, that is, ResilientDistributed Dataset, or simply RDD for short, is an abstract distributed dataset and the core abstract data type provided by the Spark framework. It can be understood as a collection of elements and can be operated in parallel in the Spark framework. The above first Resilient Distributed Dataset may include two types of values, the first type is the input range, and the second type is its corresponding output range. In one example, the input range may include the starting line number of the input, the ending line number of the input, the starting track number of the input, and the ending track number of the input. The output range may include: the starting line number of the output, the ending line number of the output, the starting track number of the output, and the ending track number of the output. The starting line number of the input and the ending line number of the input can define a range of input line numbers, the starting track number of the input and the ending track number of the input can define a range of input track numbers, the starting line number of the output and the ending line number of the output can define a range of output line numbers, and the starting track number of the output and the ending track number of the output can define a range of output track numbers. It should be understood that Step 2 can be processed through a preset function. By inputting the input range and the output range into this preset function, the first Resilient Distributed Dataset can be automatically obtained. The specific code is not limited in this embodiment of the present invention, and developers can set it by themselves.
[0040] Step 3, generate a second Resilient Distributed Dataset according to the first Resilient Distributed Dataset. Among them, the second Resilient Distributed Dataset is a key-value pair dataset obtained by calculating and reconstructing the first Resilient Distributed Dataset. The key corresponds to the line number and the track number, and the value corresponds to the single-trace data. Exemplarily, the second Resilient Distributed Dataset can be represented in the form of key-value pairs. Its key (or key value) can be (line number, track number), and its value (or value value) can be the single-trace data. The first Resilient Distributed Dataset can access the specific single-trace data. By using the calculation rules preset by the developers of this set of Spark framework (not limited in this embodiment of the present invention and can be determined according to the actual situation), the calculated single-trace data can be obtained. The second Resilient Distributed Dataset is used to represent the single-trace data after the operation.
[0041] In a possible implementation manner, step 3 may include: obtaining input data corresponding to each input range in parallel according to multiple input ranges in the first resilient distributed dataset. Exemplarily, multiple single-trace data can be directly retrieved through multiple input ranges in the first resilient distributed dataset. Since the embodiments of the present invention use the Spark framework and index data through the resilient distributed dataset, the processes of obtaining input data corresponding to each input range and the calculation process of each input data can be parallel, thereby reducing the processing time of seismic data. Then, according to the input data corresponding to each input range and the output range corresponding to each input range, an array formed by multiple single-trace data corresponding to each input data is generated in parallel. Exemplarily, after obtaining the input data, it can be calculated according to a preset calculation rule set by the developer (not limited in the embodiments of the present invention and can be determined according to actual requirements) to obtain multiple single-trace data that meet the output range. Finally, each single-trace data is separated according to the size of the byte array stored in the single-trace data and reconstructed into a key-value pair form as the second resilient distributed dataset. Among them, the key is a structure composed of a line number and a trace number, and the value is an array of single-trace data bodies. So far, the line numbers and trace numbers in the second resilient distributed dataset and the single-trace data corresponding to each line number and trace number have been calculated. It should be understood that the above steps can be processed by a preset function. By inputting the input range and the output range into the preset function, it can automatically calculate and process the input data, and the specific calculation process can be determined by the developer according to the actual situation.
[0042] Step 4, perform data integration and sorting on the second resilient distributed dataset according to the preset order between multiple line numbers and the preset order between multiple trace numbers corresponding to each line number, and use it as the processing result. Exemplarily, since the processes of obtaining the input data and the calculation process above are both parallel, the line numbers, trace numbers, and corresponding single-trace data in the obtained second resilient distributed dataset are obtained, but the order of the obtained dataset may be in a disordered state. Therefore, in order to make the data in the second resilient distributed dataset meet the requirements of subsequent seismic data processing rules, it is necessary to reorder and integrate the data in the second resilient distributed dataset. The multiple single-trace data corresponding to multiple trace numbers corresponding to the same line number are used as the trace gather corresponding to the line number, and the data integration can be achieved. The preset order between multiple line numbers and the preset order between multiple trace numbers corresponding to each line number are determined by the key-value (line number, trace number) (or called key value) in the second resilient distributed dataset. Sorting according to the key value and reconstructing it into a trace gather arranged by line number, that is, the post-stack seismic dataset.
[0043] In a possible implementation, step 4 may include: obtaining, in sequence according to a preset order among a plurality of trace numbers corresponding to each line number, the single-trace data of each trace number corresponding to each line number in the second resilient distributed dataset, and integrating the sequentially obtained plurality of single-trace data into a trace gather corresponding to the line number. Then, according to the preset order among the plurality of line numbers, sort the trace gathers corresponding to each line number, and use it as the processing result. For example: line number 1 may include trace numbers 1 to 5, and trace numbers 1 to 5 respectively correspond to single-trace data 1 to 5. Line number 2 may include trace numbers 6 to 10, and trace numbers 6 to 10 respectively correspond to single-trace data 6 to 10. Line number 3 may include trace numbers 11 to 15, and trace numbers 11 to 15 respectively correspond to single-trace data 11 to 15. If the existing line number order in the second resilient distributed dataset is 1, 3, 2, then the trace gathers 1 to 3 corresponding to line numbers 1 to 3 can be determined first, that is, trace gathers 1 to 3 are respectively single-trace data 1 to 5, single-trace data 6 to 10, and single-trace data 11 to 15. Then, sort the existing line number order of 1, 3, 2. After sorting, the line number order is 1, 2, 3, and the corresponding trace gathers are trace gathers 1, 2, 3 in sequence, that is, the integration of the processing result is realized, and the sequential arrangement and integration also meet the actual needs of developers. The processing result can also be represented as key-value pairs, where the key is the line number and the value is the trace gather corresponding to the line number.
[0044] A method for processing seismic data based on the Spark framework provided by an embodiment of the present invention can perform partitioning processing on seismic data to obtain a plurality of window range datasets, and then construct a first resilient distributed dataset according to the input range and the corresponding output range corresponding to each window range dataset. Then, based on the first resilient distributed dataset, generate a second resilient distributed dataset. Finally, according to the preset order among the plurality of line numbers and the preset order among the plurality of trace numbers corresponding to each line number, perform data integration and sorting on the second resilient distributed dataset, and use it as the processing result. The processing method provided by the embodiment of the present invention defines the window range in the first resilient distributed dataset by setting the input range and the output range to replace manual partitioning, which can reduce labor costs and labor time. In addition, in the process of generating the second resilient distributed dataset from the first resilient distributed dataset, parallel processing can be achieved relying on the Spark framework, thereby greatly reducing the processing time.
[0045] Example 2
[0046] According to an embodiment of the present invention, a processing device for seismic data based on the Spark framework is provided. Applied to the Spark framework, the processing device includes: a data partitioning module for partitioning seismic data to obtain multiple window range data sets; a first resilient distributed data set construction module for constructing a first resilient distributed data set according to the input range and the corresponding output range corresponding to each window range data set; a second resilient distributed data set construction module for generating a second resilient distributed data set according to the first resilient distributed data set; wherein, the second resilient distributed data set is a key-value pair data set obtained by computational reconstruction from the first resilient distributed data set, the key corresponds to the line number and the trace number, and the value corresponds to the single-trace data, and is used to perform data integration and sorting on the second resilient distributed data set according to the preset order between multiple line numbers and the preset order between multiple trace numbers corresponding to each line number, and use it as the processing result.
[0047] In some embodiments, the data partitioning module is further configured to obtain the input parameter size and the output parameter size; and sample the seismic data sequentially according to the input parameter size to obtain multiple window range data sets, the input range corresponding to each window range data set, and the output range corresponding to each window range data set.
[0048] In some embodiments, the second resilient distributed data set construction module is further configured to obtain the input data corresponding to each input range in parallel according to the multiple input ranges in the first resilient distributed data set; generate an array formed by multiple single-trace data corresponding to each input data in parallel according to the input data corresponding to each input range and the output range corresponding to each input range; separate each single-trace data according to the byte array size of the stored single-trace data, and reconstruct it into a key-value pair form as the second resilient distributed data set; wherein, the key is a structure composed of the line number and the trace number, and the value is an array of single-trace data bodies.
[0049] In some embodiments, the processing result generation module is further configured to sequentially obtain the single-trace data of each trace number corresponding to each line number in the second resilient distributed data set according to the preset order between multiple trace numbers corresponding to each line number, and integrate the sequentially obtained multiple single-trace data into the trace gather corresponding to the line number; sort the trace gathers corresponding to each line number according to the preset order between multiple line numbers, and use it as the processing result.
[0050] In some embodiments, the input range includes the input start line number, the input end line number, the input start trace number, and the input end trace number; the output range includes: the output start line number, the output end line number, the output start trace number, and the output end trace number.
[0051] Example 3
[0052] According to another aspect of the present invention, an electronic device is also provided. The electronic device includes:
[0053] a memory storing executable instructions:
[0054] a processor that runs the executable instructions in the memory to implement the method for processing seismic data based on the Spark framework according to the present invention.
[0055] The method includes the following steps: performing partitioning processing on seismic data to obtain a plurality of window range data sets; constructing a first resilient distributed data set according to the input range and the corresponding output range corresponding to each window range data set; generating a second resilient distributed data set according to the first resilient distributed data set; wherein, the second resilient distributed data set is a key-value pair data set obtained by computational reconstruction from the first resilient distributed data set, the key corresponds to the line number and the trace number, and the value corresponds to the single-trace data; performing data integration and sorting on the second resilient distributed data set according to the preset order between multiple line numbers and the preset order between multiple trace numbers corresponding to each line number, and using it as the processing result.
[0056] In some embodiments, the performing partitioning processing on seismic data to obtain a plurality of window range data sets includes: obtaining the input parameter size and the output parameter size; sequentially sampling the seismic data according to the input parameter size to obtain a plurality of window range data sets, the input range corresponding to each window range data set, and the output range corresponding to each window range data set.
[0057] In some embodiments, the generating a second resilient distributed data set according to the first resilient distributed data set includes: obtaining the input data corresponding to each input range in parallel according to the multiple input ranges in the first resilient distributed data set; generating an array formed by multiple single-trace data corresponding to each input data in parallel according to the input data corresponding to each input range and the output range corresponding to each input range; separating each single-trace data according to the byte array size of the single-trace data storage and reconstructing it into a key-value pair form as the second resilient distributed data set; wherein, the key is a structure composed of the line number and the trace number, and the value is an array of single-trace data bodies.
[0058] In some embodiments, integrating and sorting the second resilient distributed dataset according to a preset order among a plurality of line numbers and a preset order among a plurality of trace numbers corresponding to each line number, and using the result as a processing result, includes: sequentially obtaining single-trace data of each trace number corresponding to each line number in the second resilient distributed dataset according to the preset order among a plurality of trace numbers corresponding to each line number, and integrating the sequentially obtained multiple single-trace data into a trace gather corresponding to the line number; sorting the trace gathers corresponding to each line number according to the preset order among the plurality of line numbers, and using the result as a processing result.
[0059] In some embodiments, the input range includes an input start line number, an input end line number, an input start trace number, and an input end trace number; the output range includes: an output start line number, an output end line number, an output start trace number, and an output end trace number.
[0060] Example 4
[0061] According to another aspect of the present invention, there is also provided a computer-readable storage medium storing a computer program which, when executed by a processor, implements the method for processing seismic data based on the Spark framework according to the present invention.
[0062] The method includes the following steps: performing partitioning processing on seismic data to obtain a plurality of window range datasets; constructing a first resilient distributed dataset according to the input range and the corresponding output range corresponding to each window range dataset; generating a second resilient distributed dataset according to the first resilient distributed dataset; wherein, the second resilient distributed dataset is a key-value pair dataset obtained by computational reconstruction from the first resilient distributed dataset, the key corresponds to a line number and a trace number, and the value corresponds to single-trace data; integrating and sorting the second resilient distributed dataset according to a preset order among a plurality of line numbers and a preset order among a plurality of trace numbers corresponding to each line number, and using the result as a processing result.
[0063] In some embodiments, the performing partitioning processing on seismic data to obtain a plurality of window range datasets includes: obtaining the input parameter size and the output parameter size; sequentially sampling the seismic data according to the input parameter size to obtain a plurality of window range datasets, the input range corresponding to each window range dataset, and the output range corresponding to each window range dataset.
[0064] In some embodiments, generating a second resilient distributed dataset according to the first resilient distributed dataset includes: obtaining input data corresponding to each input range in parallel according to multiple input ranges in the first resilient distributed dataset; generating, in parallel according to the input data corresponding to each input range and the output range corresponding to each input range, an array formed by multiple single-channel data corresponding to each input data; separating each single-channel data according to the size of the byte array stored in the single-channel data and reconstructing it into a key-value pair form as the second resilient distributed dataset; wherein the key is a structure composed of a line number and a channel number, and the value is an array of single-channel data bodies.
[0065] In some embodiments, performing data integration and sorting on the second resilient distributed dataset according to a preset order between multiple line numbers and a preset order between multiple channel numbers corresponding to each line number, and using it as a processing result, includes: sequentially obtaining the single-channel data of each channel number corresponding to each line number in the second resilient distributed dataset according to the preset order between multiple channel numbers corresponding to each line number, and integrating the sequentially obtained multiple single-channel data into a channel set corresponding to the line number; sorting the channel sets corresponding to each line number according to the preset order between multiple line numbers and using it as the processing result.
[0066] In some embodiments, the input range includes an input start line number, an input end line number, an input start channel number, and an input end channel number; the output range includes: an output start line number, an output end line number, an output start channel number, and an output end channel number.
[0067] Example 5
[0068] In an embodiment of the present invention, by using Spark parallel technology, an input processing algorithm for rectangular window scanning with overlapping parts is designed. In combination with the actual scenario, first, seismic data is input in this method, and then according to the settings, the input size and output size (i.e., the input parameter size and output parameter size in the above text) are set. Then, the input line number (i.e., the input start line number and input end line number in the above text), output line number (i.e., the output start line number and output end line number in the above text), input trace number (i.e., the input start trace number and input end trace number in the above text), and output trace number (i.e., the output start trace number and output end trace number in the above text) are obtained. The entire input trace gather is divided into ranges, and the divided range windows have overlapping parts. An RDD (i.e., the first resilient distributed dataset in the above text) is constructed according to the divided ranges. Then, calculations and outputs are performed according to the input parameter range and output parameter range. Finally, the output RDD (i.e., the second resilient distributed dataset in the above text) is reconstructed, and merged processing is performed according to the same line number for output, realizing the superposition input operation processing by window scanning. The embodiment of the present invention can quickly and massively perform window scanning processing of seismic data. The manual participation part of developers is less, which is also beneficial to reducing labor costs.
[0069] For other detailed descriptions of this exemplary embodiment, reference may be made to the corresponding descriptions in the foregoing embodiments, which will not be elaborated herein.
[0070] The embodiments of the present invention have been described above. The above description is exemplary and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to the technology in the market, or to enable other ordinary skilled persons in the technical field to understand the embodiments disclosed herein.
Claims
1. A method for processing seismic data based on the Spark framework, characterized in that, applied to the Spark framework, the processing method includes: Dividing the seismic data to obtain multiple window range data sets; Constructing a first resilient distributed data set according to the input range and the corresponding output range corresponding to each window range data set; Generating a second resilient distributed data set according to the first resilient distributed data set; wherein, the second resilient distributed data set is a key-value pair data set obtained by computational reconstruction from the first resilient distributed data set, the key corresponds to the line number and the trace number, and the value corresponds to the single-trace data; Integrating and sorting the second resilient distributed data set according to the preset order between multiple line numbers and the preset order between multiple trace numbers corresponding to each line number, and using it as the processing result.
2. The processing method according to claim 1, characterized in that, The dividing and processing the seismic data to obtain multiple window range data sets includes: Obtaining the input parameter size and the output parameter size; Sampling the seismic data in sequence according to the input parameter size to obtain multiple window range data sets, the input range corresponding to each window range data set, and the output range corresponding to each window range data set.
3. The processing method according to claim 1, characterized in that, The generating the second resilient distributed data set according to the first resilient distributed data set includes: Parallelly obtaining the input data corresponding to each input range according to the multiple input ranges in the first resilient distributed data set; Parallelly generating an array formed by multiple single-trace data corresponding to each input data according to the input data corresponding to each input range and the output range corresponding to each input range; Separating each single-trace data according to the byte array size of the single-trace data storage, and reconstructing it into a key-value pair form as the second resilient distributed data set; wherein, the key is a structure composed of the line number and the trace number, and the value is an array of single-trace data bodies.
4. The processing method according to claim 1, characterized in that, The integrating and sorting the second resilient distributed data set according to the preset order between multiple line numbers and the preset order between multiple trace numbers corresponding to each line number, and using it as the processing result includes: Sequentially obtaining the single-trace data of each trace number corresponding to each line number in the second resilient distributed data set according to the preset order between multiple trace numbers corresponding to each line number, and integrating the sequentially obtained multiple single-trace data into the trace gather corresponding to the line number; Sorting the trace gathers corresponding to each line number according to the preset order between multiple line numbers, and using it as the processing result.
5. The processing method according to claim 1, characterized in that, The input range includes the input start line number, the input end line number, the input start trace number, and the input end trace number; the output range includes: the output start line number, the output end line number, the output start trace number, and the output end trace number.
6. A processing device for seismic data based on the Spark framework, characterized in that, applied to the Spark framework, the processing device includes: A data partitioning module for partitioning seismic data to obtain multiple window range data sets; A first elastic distributed data set construction module for constructing a first elastic distributed data set according to the input range and the corresponding output range corresponding to each window range data set; A second elastic distributed data set construction module for generating a second elastic distributed data set according to the first elastic distributed data set; wherein, the second elastic distributed data set is a key-value pair data set obtained by computational reconstruction from the first elastic distributed data set, the key corresponds to the line number and the trace number, and the value corresponds to the single-trace data; A processing result generation module for integrating and sorting the second elastic distributed data set according to a preset order between multiple line numbers and a preset order between multiple trace numbers corresponding to each line number, and using it as a processing result.
7. The processing device according to claim 6, wherein, the data partitioning module is further configured to obtain the input parameter size and the output parameter size; and sample the seismic data in sequence according to the input parameter size to obtain multiple window range data sets, the input range corresponding to each window range data set, and the output range corresponding to each window range data set.
8. The processing device according to claim 6, wherein, the second elastic distributed data set construction module is further configured to obtain the input data corresponding to each input range in parallel according to the multiple input ranges in the first elastic distributed data set; and generate an array formed by multiple single-trace data corresponding to each input data in parallel according to the input data corresponding to each input range and the output range corresponding to each input range; separate each single-trace data according to the byte array size of the single-trace data storage, and reconstruct it into a key-value pair form as the second elastic distributed data set; wherein, the key is a structure composed of the line number and the trace number, and the value is an array of single-trace data bodies.
9. The processing device according to claim 6, wherein, the processing result generation module is further configured to sequentially obtain the single-trace data of each trace number corresponding to each line number in the second elastic distributed data set according to the preset order between the multiple trace numbers corresponding to each line number, and integrate the sequentially obtained multiple single-trace data into the trace gather corresponding to the line number; and sort the trace gathers corresponding to each line number according to the preset order between multiple line numbers, and use it as a processing result.
10. The processing device according to claim 6, wherein, the input range includes the input start line number, the input end line number, the input start trace number, and the input end trace number; the output range includes: the output start line number, the output end line number, the output start trace number, and the output end trace number.
11. An electronic device, wherein, the electronic device includes: a memory storing executable instructions; a processor that runs the executable instructions in the memory to implement the processing method according to any one of claims 1-5.
12. A computer-readable storage medium storing a computer program which, when executed by a processor, implements the processing method according to any one of claims 1-5.