Seismic operation execution method, device and system
By allocating sub-jobs of seismic data between the main working node and the slave working node and storing data in the shared disk, parallel processing of seismic data is realized, solving the problem of slow processing speed of large-scale seismic data in the prior art, and improving processing efficiency and stability.
Patent Information
- Application Number
- CN202311566828.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-22
- Publication Date
- 2025-05-23
AI Technical Summary
The prior art is slow when processing large-scale seismic data, and when the program process is stuck, the data processing speed is further reduced, which cannot meet the needs of large-scale seismic data processing.
Parallel processing is achieved by allocating sub-jobs between the primary and slave job nodes and storing seismic data in the shared disk. The main working node divides the operations into the smallest units based on the seismic track set, and allocates sub-jobs according to the processing capabilities of the secondary working nodes, thereby improving processing efficiency.
Through parallel processing, the processing speed and stability of large-scale seismic data is significantly improved, and the program process is avoided.
Smart Images

Figure CN120029724A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the technical field of seismic data processing, and in particular to a method, device and system for executing seismic operations. Background Art
[0002] With the rapid growth of seismic data collection, the serial seismic operation execution mode is slow in processing seismic data and can no longer meet the needs of large-scale seismic data processing. In the case of program process jamming, the seismic data processing speed is further reduced. Summary of the invention
[0003] The present application provides a method, device and system for executing seismic operations. The technical solution is as follows:
[0004] On the one hand, an embodiment of the present application provides a method for performing seismic operations, the method comprising:
[0005] The master operation node divides the sub-operation into sub-operations according to the number of sub-operation divisions of the seismic data in the master operation, taking the seismic gather as the minimum division unit, and obtains the sub-operation information of each sub-operation, wherein the seismic data is stored in a shared disk, and the master operation node and each slave operation node support access to the shared disk;
[0006] The master operation node allocates the sub-operation to the slave operation node according to the node processing capability of each slave operation node;
[0007] The slave operation node extracts sub-operation seismic data from the seismic data according to the sub-operation information of the sub-operation, and performs data processing on the sub-operation seismic data to obtain a sub-operation processing result;
[0008] The slave job node writes the sub-job processing result into the shared disk;
[0009] When all sub-jobs are processed, the main operation node merges the sub-operation processing results of each sub-job to obtain the main operation data processing result of the seismic data.
[0010] On the other hand, an embodiment of the present application provides a seismic operation execution device, the device comprising:
[0011] A main operation module, used to divide the sub-operation into sub-operations according to the number of sub-operation divisions of the seismic data in the main operation, taking the seismic gather as the minimum division unit, and obtain the sub-operation information of each sub-operation, wherein the seismic data is stored in a shared disk, and the main operation module and each slave operation module support access to the shared disk;
[0012] The master operation module is used to allocate the sub-operations to the slave operation modules according to the node processing capabilities of the slave operation modules;
[0013] The sub-job module is used to extract sub-job seismic data from the seismic data according to the sub-job information of the sub-job, and perform data processing on the sub-job seismic data to obtain a sub-job processing result;
[0014] The slave job module is used to write the sub-job processing result into the shared disk;
[0015] The main operation module is used to merge the sub-operation processing results of each sub-operation when all sub-operations are processed, so as to obtain the main operation data processing result of the seismic data.
[0016] On the other hand, an embodiment of the present application provides a computer system, the computer system comprising a master operation node and a plurality of slave operation nodes;
[0017] The master operation node is used to divide the sub-operations according to the number of sub-operations of the seismic data in the master operation, taking the seismic gather as the minimum division unit, and obtain the sub-operation information of each sub-operation. The seismic data is stored in a shared disk, and the master operation node and each of the slave operation nodes support access to the shared disk;
[0018] The master operation node is used to allocate the sub-operation to the slave operation node according to the node processing capability of each slave operation node;
[0019] The slave operation node is used to extract sub-operation seismic data from the seismic data according to the sub-operation information of the sub-operation, and perform data processing on the sub-operation seismic data to obtain a sub-operation processing result;
[0020] The slave job node is used to write the sub-job processing result into the shared disk;
[0021] The main operation node is used to merge the sub-operation processing results of each sub-operation when all sub-operations are processed, so as to obtain the main operation data processing result of the seismic data.
[0022] On the other hand, an embodiment of the present application provides a computer device, the computer device comprising a processor and a memory, the memory storing at least one program code, the at least one program code being loaded and executed by the processor to implement the seismic operation execution method as described in the above aspects. The computer device can be implemented as the master operation node or the slave operation node as described in the above aspects.
[0023] On the other hand, an embodiment of the present application provides a computer-readable storage medium, wherein the storage medium stores at least one program code, and the at least one program code is used to be executed by a processor to implement the seismic operation execution method as described in the above aspects.
[0024] In an embodiment of the present application, the master operation node divides the sub-operations into smaller units based on the seismic track set, and assigns the sub-operations to each slave operation node. The slave operation node then processes the seismic data of the sub-operations obtained by the division, and stores the processing results into a shared disk. Finally, the master operation node merges the processing results of each sub-operation to obtain the main operation data processing result, wherein each sub-operation runs in parallel on each slave operation node, thereby improving the processing speed and stability of large-scale seismic data. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0026] Figure 1 A schematic diagram showing an implementation environment provided by an exemplary embodiment of the present application is shown;
[0027] Figure 2 A flow chart showing a method for performing seismic operations provided by an exemplary embodiment of the present application is shown;
[0028] Figure 3 A schematic diagram of an implementation of a seismic operation execution process provided by an exemplary embodiment of the present application is shown;
[0029] Figure 4 A flow chart showing a method for performing seismic operations provided by another exemplary embodiment of the present application is shown;
[0030] Figure 5 is a schematic diagram of the division of seismic gathers provided by an exemplary embodiment of the present application;
[0031] Figure 6 is a flowchart of a main job restart process provided by an exemplary embodiment of the present application;
[0032] Figure 7 It is a schematic diagram of implementing the change of the sub-job status provided by an exemplary embodiment of the present application;
[0033] Figure 8 is a schematic diagram of an interactive interface for executing seismic operations provided by an exemplary embodiment of the present application;
[0034] Fig. 9 A structural block diagram of a seismic operation execution device provided by an exemplary embodiment of the present application is shown;
[0035] Fig.10 A computer system structure diagram provided by an exemplary embodiment of the present application is shown;
[0036] Fig.11 A structural block diagram of a computer device provided by an exemplary embodiment of the present application is shown. DETAILED DESCRIPTION
[0037] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.
[0038] Please refer to Figure 1 , which shows a schematic diagram of an implementation environment provided by an exemplary embodiment of the present application. The implementation environment includes a master operation node 110, a slave operation node 120 and a shared disk 130.
[0039] The master operation node 110 is a computing node for executing the master operation, which is mainly used for sub-operation allocation, seismic data division, and processing result merging. The slave operation node 120 is a computing node for executing sub-operations, which is mainly used for seismic data processing and operation and maintenance. The shared disk 130 is used to store seismic data. The specific work and execution order of the master operation node 110 and several slave operation nodes 120 are as follows.
[0040] The main operation node 110 is used to divide the sub-operations according to the number of sub-operations of the seismic data in the main operation, taking the seismic track set as the minimum division unit, and obtain the sub-operation information of each sub-operation. The seismic data is stored in the shared disk 130. The main operation node 110 and each slave operation node 120 support access to the shared disk 130.
[0041] The master operation node 110 is further used to allocate sub-operations to the slave operation nodes 120 according to the node processing capabilities of the slave operation nodes 120 .
[0042] The slave operation node 120 is used to extract sub-operation seismic data from the seismic data according to the sub-operation information of the sub-operation, and perform data processing on the sub-operation seismic data to obtain the sub-operation processing result.
[0043] The slave job node 120 is used to write the sub-job processing result into the shared disk 130 .
[0044] The main operation node 110 is also used to merge the sub-operation processing results of each sub-operation when all sub-operations are processed, so as to obtain the main operation data processing result of the seismic data.
[0045] The solution provided in the embodiment of the present application can be run on multiple hardware devices, such as computers, servers, etc., or can be run on a single computer device, and functions such as job allocation and data processing can be realized through the main process and sub-processes. The solution is described in detail below based on this implementation environment.
[0046] Please refer to Figure 2 , which shows a flowchart of a method for performing seismic operations provided by an exemplary embodiment of the present application, the method may include the following steps:
[0047] Step 201, the main operation node divides the sub-operations according to the number of sub-operations of the seismic data in the main operation, taking the seismic track set as the minimum division unit, and obtains the sub-operation information of each sub-operation. The seismic data is stored in the shared disk, and the main operation node and each slave operation node support access to the shared disk.
[0048] The main job is a complete seismic job for seismic data processing, which includes a program for implementing the seismic data processing function. The main job node divides the main job into sub-jobs so that each sub-job implements the seismic data processing function of the main job.
[0049] The number of sub-job divisions is a predefined command line parameter, which is used to limit the total number of sub-jobs obtained by division. In one possible implementation, when the process of the main job node is started, the determined number of sub-job divisions is passed to the process of the main job node as a command line parameter to perform sub-job division.
[0050] The seismic trace includes a detector, an amplification system and a recording system, which is used to record seismic reflection waves or refracted waves propagating to the seismic measuring point.
[0051] Seismic trace data consists of a series of seismic waveform traces.
[0052] Seismic data consists of a series of seismic traces with fixed length.
[0053] A seismic track gather is a collection of multiple seismic tracks with common attributes, for example, a common shot point track gather, which is a track gather consisting of all seismic tracks excited by the same shot point and received by different detection points, where the shot point (excitation point) is the location where the shot is fired, and the detection point (receiving point) is the location where the seismic waves are received.
[0054] The sub-job information is used to describe the relevant information of each sub-job obtained by division, such as the job name, job number, and the scope of seismic data processed by each sub-job.
[0055] The shared disk is used to store the original seismic data and the results of seismic data processing. The master job node and each slave job node support access to the shared disk. In a possible implementation, the master job node divides the sub-jobs according to the data range of the seismic data in the main job and the position in the shared disk, obtains the data range of the seismic data in the sub-jobs and the position in the shared disk, then obtains the seismic data of the sub-jobs from the shared disk from the slave job node for data processing, and finally stores the data processing results of the sub-jobs in the shared disk from the slave job node.
[0056] Regarding the sub-job division process, in a possible implementation, the main job node first determines the unit of the seismic track gather, for example, the seismic track gather is determined by shot, the seismic data in the main job is composed of multiple seismic track gathers, and the main job node divides the sub-jobs according to the number of sub-job divisions, wherein one sub-job contains one or more seismic track gathers, and multiple seismic track gathers are arranged continuously in the seismic data.
[0057] Step 202: The master operation node allocates sub-operations to the slave operation nodes according to the node processing capabilities of the slave operation nodes.
[0058] The node processing capacity is used to characterize the number of sub-jobs that each slave job node can process simultaneously, that is, the maximum number of sub-jobs that can be processed.
[0059] Optionally, the master job node may extract the maximum processing quantity of sub-jobs from a file describing the node processing capabilities of each slave job node, thereby allocating sub-jobs to each slave job node.
[0060] Step 203, extracting sub-job seismic data from the seismic data from the operation node according to the sub-job information of the sub-job, and performing data processing on the sub-job seismic data to obtain a sub-job processing result.
[0061] In one possible implementation, the sub-job performs the same processing on the assigned seismic data according to the main job's processing of the seismic data, wherein the main job's processing of the seismic data is a predetermined content, and the specific processing process is determined based on actual seismic data processing needs.
[0062] Step 204: Write the sub-job processing result from the job node to the shared disk.
[0063] In a possible implementation, the slave job node writes the sub-job processing results to a location in a shared disk for storing seismic data processing results, such as a processing result folder, so that the master job node can subsequently merge the sub-job processing results of each sub-job.
[0064] Step 205: When all sub-jobs are processed, the main operation node merges the sub-operation processing results of each sub-operation to obtain the main operation data processing result of the seismic data.
[0065] Since the sub-job processes seismic data in units of seismic gathers, the integrity of the seismic data in the sub-job is guaranteed. Therefore, the division and merging of seismic data by the main job node will not affect the main job data processing results.
[0066] In a possible implementation, the main operation node merges the sub-operation processing results of each sub-operation according to the arrangement order of the original seismic data to obtain continuous main operation data processing results.
[0067] For example, Figure 3 The schematic diagram of the implementation of the seismic operation execution process provided by an exemplary embodiment of the present application is shown. The main operation node 301 first performs seismic data segmentation to obtain sub-operation seismic data 302, sub-operation seismic data 303 and sub-operation seismic data 304, then the slave operation node 305 processes the sub-operation seismic data 302 to obtain the sub-operation processing result 308, the slave operation node 306 processes the sub-operation seismic data 303 to obtain the sub-operation processing result 309, and the slave operation node 307 processes the sub-operation seismic data 304 to obtain the sub-operation processing result 310, wherein the sub-operation seismic data 302, the sub-operation seismic data 303 and the sub-operation seismic data 304 are obtained by segmenting the seismic data 311. Finally, the main operation node 301 merges the sub-operation processing results 308, the sub-operation processing results 309 and the sub-operation processing results 310 to obtain the main operation data processing result 312.
[0068] To summarize, in the embodiment of the present application, the master operation node divides the sub-operations into smaller units based on the seismic track set, and assigns the sub-operations to each slave operation node. The slave operation node then processes the seismic data of the sub-operations obtained by the division, and stores the processing results into a shared disk. Finally, the master operation node merges the processing results of each sub-operation to obtain the main operation data processing result, wherein each sub-operation runs in parallel on each slave operation node, thereby improving the processing speed and stability of large-scale seismic data.
[0069] Please refer to Figure 4 , which shows a flow chart of a method for performing seismic operations provided by another exemplary embodiment of the present application, the method may include the following steps:
[0070] Step 401: The main operation node determines the total number of seismic trace data according to the index keyword of the seismic data and the main operation index keyword range.
[0071] The index keyword and the main operation index keyword range are the operation parameters of the main operation, which are used to describe the data range of the seismic data processed by the main operation. Among them, the index keyword is used to represent the hierarchical type of the data processed by the main operation, such as the shot number, track number, line number, shot point coordinates, etc. The index keyword range is used to represent the continuous range of seismic data described by the index keyword.
[0072] Exemplarily, when the index keyword is the shot number, the index keyword range can be expressed as 1 to 10,000, indicating that the range of seismic data processed by the main job is seismic data with shot numbers from 1 to 10,000.
[0073] In some possible implementations, the main job node obtains the index keyword of the seismic data and the main job index keyword range from the file describing the seismic data processing rules in the main job.
[0074] Regarding the determination of the total number of seismic channel data, in one possible implementation, the main operation node counts the seismic channel data that fall within the description range of the index keyword and the index keyword range in the index table based on the index keyword of the seismic data and the main operation index keyword range, and obtains the total number of seismic channel data and the number of seismic channel data contained in each seismic channel set.
[0075] Exemplarily, the index table is shown in Table 1, wherein the channel number is used to indicate the number of each seismic channel data, and the index keyword 1 and the index keyword 2 are used to index and locate the seismic channel data.
[0076] It should be noted that the number of index keywords in the index table is not fixed, that is, the index table can use multi-level index keywords to locate the seismic channel data. For example, using the secondary index keywords, that is, the shot number and the channel number, when the shot number is 1 and the channel number is 100, the offset of the seismic channel data of the 100th channel of the 1st shot can be determined, so as to extract the seismic channel data of the 100th channel. Among them, the offset is used to characterize the specific byte position of the seismic channel data in the seismic data file.
[0077] Table 1
[0078] road number Index keyword 1 Index keyword 2 Offset 1 1 1 0 2 1 1 10 3 1 1 20 4 1 1 30 …… …… …… …… 101 1 1 1010 102 1 2 1020 …… …… …… …… 230 1 2 2300 231 2 1 2310 232 2 1 2320 …… …… …… …… p n m q
[0079] The index keyword of seismic data is index keyword 1, and the main operation index keyword range is 1 to n. Therefore, the total number of seismic channel data is p, and the pth seismic channel data corresponds to index keyword 2 numbered m, and the offset is q, where n, p, m and q are positive integers.
[0080] This example divides the seismic channel set according to the number of index keyword 2, that is, the continuous seismic channel data with the same index keyword 2 number are divided into one seismic channel set. For example, in Table 1, the seismic channel data with channel numbers from 1 to 101 are divided into the first seismic channel set, and the seismic channel data with channel numbers from 102 to 230 are divided into the second seismic channel set. Therefore, the first seismic channel set contains 101 seismic channel data, and the second seismic channel set contains 129 seismic channel data.
[0081] Step 402: The main operation node determines the number of reference channels of seismic channel data of the sub-operation according to the total number of seismic channel data and the number of sub-operation divisions.
[0082] The reference number of seismic data of the sub-job is used to measure the number of seismic data in each sub-job. That is, after the main job node is allocated, the number of seismic data in the sub-job is close to the reference number of seismic data of the sub-job, thereby ensuring that the difference in the number of seismic data allocated to each sub-job is small.
[0083] In a possible implementation, the main operation node calculates the average number of seismic trace data allocated to each sub-operation according to the total number of seismic trace data and the number of sub-operation divisions, and then determines the average value as the reference number of seismic trace data for the sub-operation.
[0084] Step 403, the main operation node divides the sub-operations according to the number of seismic data reference channels and the number of seismic data channels contained in each seismic channel gather, taking the seismic channel gather as the minimum division unit, and obtains the sub-operation information of each sub-operation.
[0085] Optionally, the sub-job information of each sub-job may include an index keyword of each sub-job and a sub-job index keyword range.
[0086] In one possible implementation, the main job node divides the seismic track sets with consecutive numbers into a sub-job so that the sum of the number of seismic data channels in the seismic track sets with consecutive numbers is close to the number of reference seismic data channels to a certain extent, thereby determining the sub-job index keyword range of the sub-job information according to the seismic track set range.
[0087] The specific division of the seismic gathers can be determined according to the threshold, including the following two cases, where i and j are positive integers.
[0088] Case 1: When the sum of the number of first seismic channel data from the i-th seismic channel set to the i+j-th seismic channel set is less than the number of reference seismic channel data, the sum of the number of second seismic channel data from the i-th seismic channel set to the i+j+1-th seismic channel set is greater than the number of reference seismic channel data, and the difference between the sum of the second seismic channel data and the number of reference seismic channel data is greater than a threshold, the main job node determines the sub-job index keyword range in the sub-job information according to the i-th seismic channel set to the i+j-th seismic channel set.
[0089] In the second case, when the sum of the number of first seismic channel data from the i-th seismic channel set to the i+j-th seismic channel set is less than the number of reference seismic channel data, the sum of the number of second seismic channel data from the i-th seismic channel set to the i+j+1-th seismic channel set is greater than the number of reference seismic channel data, and the difference between the sum of the second seismic channel data and the number of reference seismic channel data is less than the threshold, the main operation node determines the sub-operation index keyword range in the sub-operation information according to the i-th seismic channel set to the i+j+1-th seismic channel set.
[0090] For example, the division diagram of seismic gathers is as follows: Figure 5 As shown, the 1st to 99th seismic trace gathers belong to sub-job 1, and the 100th seismic trace gathers and onwards belong to sub-job 2. When the sum of the number of seismic trace data channels from the 100th to the 201st seismic trace gathers, that is, the sum of the first seismic trace data channel number, is less than the number of reference seismic trace data channels, and the sum of the number of seismic trace data channels from the 100th to the 202nd seismic trace gathers, that is, the sum of the second seismic trace data channel number, is greater than the number of reference seismic trace data channels, the main operation node calculates the difference between the sum of the second seismic trace data channel number and the number of reference seismic trace data channels, that is, the difference between the sum of the seismic trace data channel number from the 100th to the 202nd seismic trace gathers and the number of reference seismic trace data channels.
[0091] When the difference is greater than the threshold, the main job node determines the 100th seismic track gather to the 201st seismic track gather as the sub-job index keyword range of sub-job 2; when the difference is less than the threshold, the main job node determines the 100th seismic track gather to the 202nd seismic track gather as the sub-job index keyword range of sub-job 2, wherein the sub-job index keyword in the sub-job information is the index keyword corresponding to the seismic track gather.
[0092] The threshold value may be determined by the master operation node according to the difference between the sum of the second seismic data channel numbers and the reference seismic data channel numbers. For example, the average value of the difference between the sum of the second seismic data channel numbers and the reference seismic data channel numbers is calculated and the average value is determined as the threshold value.
[0093] In addition, when the sum of the number of first seismic channel data from the i-th seismic channel set to the i+j-th seismic channel set is equal to the number of reference seismic channel data, the main operation node determines the sub-job index keyword range in the sub-job information based on the i-th seismic channel set to the i+j-th seismic channel set.
[0094] Step 404: the master operation node allocates sub-operations to the slave operation nodes according to the node processing capabilities of the slave operation nodes.
[0095] Optionally, the master operation node may allocate sub-operations to the slave operation nodes according to the maximum number of sub-operations of each slave operation node recorded in the slave operation node list. This step may include the following sub-steps.
[0096] Step 404A, the master operation node obtains a list of slave operation nodes from the shared disk. The list of slave operation nodes includes the maximum number of sub-operations of each slave operation node. The maximum number of sub-operations refers to the upper limit of the number of sub-operations that can be processed in parallel by the slave operation node.
[0097] The slave job node list may include the name of each slave job node and the corresponding maximum number of sub-jobs.
[0098] In a possible implementation, the master operation node may obtain the slave operation node list from the shared disk by parsing command line parameters.
[0099] Step 404B: the master operation node allocates sub-operations to each slave operation node based on the maximum number of sub-operations of each slave operation node.
[0100] In a possible implementation, in order to ensure load balancing among slave job nodes, the master job node allocates sub-jobs to slave job nodes with a large number of assignable sub-jobs, where the number of assignable sub-jobs is the difference between the maximum number of sub-jobs and the number of sub-jobs currently being processed.
[0101] For example, the maximum number of sub-jobs of slave job node 1 is 3, the number of sub-jobs currently being processed is 2, and the number of sub-jobs that can be allocated is 1; the maximum number of sub-jobs of slave job node 2 is 5, the number of sub-jobs currently being processed is 2, and the number of sub-jobs that can be allocated is 3. Since the number of sub-jobs that can be allocated of slave job node 2 is greater than the number of sub-jobs that can be allocated of slave job node 1, the master job node allocates the sub-job currently to be allocated to slave job node 2.
[0102] Step 405, extracting sub-job seismic data from the seismic data from the operation node according to the sub-job information of the sub-job, and performing data processing on the sub-job seismic data to obtain a sub-job processing result.
[0103] In a possible implementation, the master operation node first remotely logs into the slave operation node, and then starts the sub-operation by calling the seismic operation execution control program, so that the slave operation node completes data processing of the sub-operation seismic data and obtains the sub-operation processing result.
[0104] In step 403, the main operation node assigns the seismic gathers with consecutive numbers to the sub-operations, and obtains the sub-operation index keyword range in the sub-operation information. Since the seismic gathers are divided according to the index keywords, the sub-operation information includes the index keywords describing the seismic gather division rules and the corresponding index keyword range. The sub-operation information characterizes the range of the continuous seismic gathers assigned to the sub-operation through the index keywords and the index keyword range, so as to extract the sub-operation seismic data.
[0105] From the above reasoning, it can be known that the sub-job seismic data can be extracted from the seismic data from the job node according to the index keyword contained in the sub-job information and the sub-job index keyword range.
[0106] In one possible implementation, the job node extracts index keywords and sub-job index keyword ranges from the sub-job information, then determines the storage location (offset) of the sub-job seismic data by querying the index table, and then extracts the sub-job seismic data from the seismic data file on the shared disk.
[0107] The shared disk stores a sub-job status file, which contains the running status of each sub-job, including the waiting state, running state, running success state, running failure state and canceled running state.
[0108] Optionally, the slave job node may modify the status of each sub-job in the sub-job status file during the sub-job allocation process, the running process, and the end of the running process, thereby processing the sub-job.
[0109] The waiting-to-run state is used to represent the sub-jobs to be assigned, and the master job node assigns the sub-jobs in the waiting-to-run state to the slave job nodes.
[0110] The Running state is used to indicate that the subjob is currently running.
[0111] The successful operation status is used to indicate that the sub-job operation is completed and the sub-job processing result is obtained.
[0112] The run failure status is used to indicate that the sub-job has ended due to a failure and no sub-job processing result has been generated.
[0113] The canceled operation status is used to indicate that the sub-job operation is terminated and no sub-job processing result is obtained.
[0114] It should be noted that, since each sub-job needs to modify the sub-job status file during the parallel operation of the sub-job, if multiple sub-jobs share the same sub-job status file, the sub-job status file content may be disordered due to simultaneous modification. Therefore, multiple sub-job status files can be stored in the shared disk to record the running status of each sub-job, for example, each sub-job corresponds to a sub-job status file.
[0115] Optionally, the master job node allocates sub-jobs in a waiting-to-run state to the slave job nodes according to the node processing capabilities of the slave job nodes.
[0116] In one possible implementation, the master job node first creates a sub-job based on the sub-job information, and the status of the sub-job is a waiting-to-run status. The sub-job in the waiting-to-run status is then placed in a buffer space, and then the sub-job is assigned to a slave job node for processing based on the node processing capabilities of each slave job node.
[0117] Optionally, after the slave job node is assigned to the sub-job, the slave job node starts the sub-job to process the seismic data of the sub-job, and the slave job node switches the running state of the sub-job from the waiting state to the running state.
[0118] Step 406: Write the sub-job processing result from the job node to the shared disk.
[0119] The implementation of this step may refer to the above step 204, and this embodiment will not be described in detail here.
[0120] Step 407: When all sub-jobs are processed, the main job node merges the sub-job processing results of each sub-job to obtain the main job data processing result of the seismic data.
[0121] The implementation of this step may refer to the above step 205, and this embodiment will not be described in detail here.
[0122] In the embodiment of the present application, in order to ensure load balancing, the main operation node first determines the number of reference channels of seismic data, and then divides the sub-operations according to the number of reference channels of seismic data, so that the number of seismic data in each sub-operation is similar, thereby obtaining the sub-operation information of each sub-operation, and finally the slave operation node performs data processing according to the sub-operation information. During the data processing process, the slave operation node modifies the running state of the sub-operation, so as to perform different processing on the sub-operations in different states, thereby reducing the situation where individual sub-operations are stuck during the parallel operation of sub-operations, and improving the processing speed and stability of large-scale seismic data.
[0123] After the slave job node writes the sub-job processing result to the shared disk, the sub-job successfully completes because the main job data processing result of the seismic data is obtained, and the slave job node switches the sub-job running status from the running status to the successful running status.
[0124] However, not all sub-jobs can run successfully. When running a sub-job from a job node, the sub-job may fail due to equipment failure or improper job parameter settings. The following describes the sub-job processing process from a job node when a sub-job fails.
[0125] In the case of a sub-job processing failure, the slave job node can update the sub-job failure count in the sub-job status file, and determine the sub-job running status through the sub-job failure count. Therefore, sub-job processing failure can also be divided into the following two situations.
[0126] In case 1, when the number of failures reaches the number threshold, the slave job node switches the running state of the sub-job from the running state to the running failure state.
[0127] When the number of failures reaches the threshold, for example, more than 3 times, the sub-job may fail due to non-program reasons and needs to be handled. The slave job node switches the sub-job's running state from the running state to the failed state so that manual review can be introduced after the main job is finished. After troubleshooting, the slave job node switches the sub-job's running state from the failed state to the waiting state, and restarts the main job to start the sub-job in the waiting state.
[0128] In case 2, when the number of failures does not reach the number threshold, the slave job node switches the running state of the sub-job from the running state to the waiting state.
[0129] During the sub-job processing, the sub-job may fail to run because the slave job node is busy. By restarting the sub-job, the sub-job processing result can still be obtained normally. Therefore, the slave job node can automatically restart the sub-job whose failure times do not reach the number threshold, that is, the slave job node switches the running state of the sub-job from the running state to the waiting state, and the master job node reallocates the sub-job node in the waiting state.
[0130] Since the main job is customized by the user, this application modifies the data processing scope based on the main job set by the user, thereby dividing the sub-jobs and realizing parallel processing of seismic data. Therefore, a large number of sub-jobs may fail to run due to unreasonable settings of the main job, such as unreasonable settings of the job parameters and command line parameters of the main job.
[0131] In order to solve the above problem, the main job node can determine whether the current main job needs to be restarted by setting a percentage threshold of sub-job failures. As shown in Figure 6, the main job restart process may include the following steps.
[0132] Step 601: The main job node obtains the sub-job failure number of the sub-jobs in the failed running state from the sub-job status file.
[0133] In a possible implementation, the main job node counts the number of sub-jobs in a failed state in real time, that is, when the running state of a sub-job switches from a running state to a failed state, the main job node accumulates the current number of sub-job failures.
[0134] Because the slave job node automatically switches the running state of the sub-job from the running state to the waiting state when the number of failures does not reach the number threshold, the number of sub-job failures records the sub-jobs whose number of failures reaches the number threshold and whose running fails.
[0135] Step 602: When the ratio of the number of failed sub-jobs to the total number of sub-jobs reaches a ratio threshold, the main job node terminates the running main job and restarts the main job.
[0136] Exemplarily, 20% may be taken as the ratio threshold of the number of sub-job failures to the total number of sub-jobs. When the ratio of the number of sub-job failures to the total number of sub-jobs exceeds 20%, the main job node terminates the running main job.
[0137] It should be noted that when the ratio of the number of sub-job failures to the total number of sub-jobs reaches the ratio threshold, it means that there is a problem with the main job, so the main job is terminated. After the problem of the main job is solved, the main job is restarted.
[0138] Step 603: When the main job is terminated, the slave job node switches the sub-job in the running state to the canceled state.
[0139] In one possible implementation, when the main job terminates, the main job node sends a termination instruction to the slave job node, where the termination instruction is used to indicate that the main job has stopped running. Upon receiving the termination instruction, the slave job node switches the sub-job node in the running state to the canceled running state.
[0140] Step 604: When the main job is restarted, the slave job node switches the sub-jobs in the failed operation state and the canceled operation state to the waiting state.
[0141] In a possible implementation, when the main job is restarted, the slave job node switches the sub-job from the failed operation state and the canceled operation state to the waiting state, and then the main job node assigns the sub-job in the waiting state to the slave job node for re-execution.
[0142] For example, the implementation diagram of the sub-job status change is as follows: Figure 7 As shown. After the main job node allocates the sub-job, the slave job node starts the sub-job, and the sub-job switches from the waiting state to the running state. When the sub-job runs successfully and obtains the sub-job processing result, the sub-job switches from the running state to the successful state, and the sub-job ends. When the sub-job fails, the sub-job switches from the running state to the failed state. Among them, if the number of sub-job failures does not reach the number threshold, the sub-job switches from the running state to the waiting state and re-participates in the sub-job allocation; if the number of sub-job failures reaches the number threshold, the sub-job switches to the failed state until the main job ends. During the operation of the main job, if the proportion of the number of sub-job failures to the total number of sub-jobs does not reach the proportion threshold, the main job keeps running until all sub-jobs enter the successful state. If the proportion reaches the proportion threshold, the main job terminates, and the slave job node switches the sub-jobs in the running state to the canceled state. Then, when the main job is restarted, the slave job node switches the sub-jobs in the failed state and the canceled state to the waiting state, waiting for restart.
[0143] In addition, since the slave operation node may have a low processing success rate due to equipment failure and other problems, in order to improve the processing speed of large-scale seismic data, the master operation node can record the success rate of the slave operation node through the slave operation node work log, thereby excluding the slave operation node with too low success rate. The method may include the following two steps.
[0144] Step 1: The master job node obtains the work log of the slave job node from the shared disk. The work log of the slave job node includes the success rate of each slave job node in processing the sub-job.
[0145] In one possible implementation, after the sub-job processing is completed, the main job node counts the number of sub-job failures and successes, and then calculates the ratio of the number of successes to the number of sub-job processing times, that is, the success rate, and then records the success rate in the work log of the slave job node.
[0146] Step 2: The master job node filters out slave job nodes with low success rates based on the work logs of the slave job nodes. The success rate of sub-jobs processed by the slave job nodes with low success rates is lower than the success rate threshold.
[0147] In a possible implementation, the master job node first determines a success rate threshold, and then obtains the working logs of the slave job nodes. Based on the success rates of the slave job nodes recorded in the working logs of the slave job nodes, sub-jobs are not assigned to the slave job nodes with low success rates, and the sub-jobs are run on the slave job nodes with higher success rates, thereby improving the processing speed of the sub-jobs and avoiding consuming extra time due to the failure of the sub-jobs to run.
[0148] Since restarting the master job does not require restarting the sub-jobs in the running success state, while the newly created master job needs to start running all the sub-jobs in the pending running state, in order to distinguish the restarted master job, the computer device can be identified by using the master job identifier.
[0149] In a possible implementation, the master job node first obtains the master job identifier of the master job, such as the job number of the master job, and then determines that the current master job is a restarted master job based on the existence of the master job identifier, including the following two cases.
[0150] Case 1, when the working directory corresponding to the master job identifier does not exist in the shared disk, the master job node creates a working directory for the master job, and the working directory is used to store files related to the master job and sub-jobs.
[0151] The working directory may include master job files, sub-job files, sub-job information files, slave job node lists, sub-job status files, slave job node working logs, parameter configuration files, etc.
[0152] The master job file and the sub-job file contain executable programs. The sub-job file is generated based on the master job. The executable program may include the following sub-programs: data input program, data processing program, and data output program, where the seismic data processing process is executed by the data processing program.
[0153] In some possible implementations, each slave job node realizes the seismic data processing process by running the sub-job file in the corresponding working directory on the shared disk.
[0154] The sub-job information file is used to store the sub-job information of each sub-job obtained by the master job node through partitioning.
[0155] The slave job node list includes the node names and node processing capabilities of each job node.
[0156] The sub-job status file can record the current processing status of each sub-job respectively.
[0157] The slave job node working log is used to record the log information of each slave job node.
[0158] The parameter configuration file contains some parameter configuration information related to job running, such as command line parameters.
[0159] Case 2: When a working directory corresponding to the main job identifier exists in the shared disk, the main job node continues to use the working directory.
[0160] When there is a working directory corresponding to the main job identifier in the shared disk, the main job node can determine that the current main job is a restart job, and then restart the main job based on the original working directory, including restarting all sub-jobs that are not in a successful running state, setting all sub-jobs that are not in a successful running state to a waiting-to-run state on the slave job node, and the main job node reallocates the sub-jobs in a waiting-to-run state to the slave job node for running.
[0161] It should be further explained that the main job identifier, sub-job node list, sub-job division quantity, number threshold, and ratio threshold mentioned in this application can be obtained by the main job node by parsing command line parameters when the main job is started.
[0162] In order to visualize the execution process of the seismic operation, the main operation node can display the running status and log information of the sub-jobs on the interactive interface, and set restart buttons and termination buttons for the sub-jobs and the main job.
[0163] For example, Figure 8 The figure shows a schematic diagram of the interactive interface for executing seismic operations. The interactive interface uses a list to display the sub-job number, sub-job operation status, slave job node name, number of failures, start time, and end time corresponding to each sub-job. Users can restart or terminate the main job or sub-job based on the relevant information of the current sub-job.
[0164] Please refer to Fig. 9 , which shows a structural block diagram of a seismic operation execution device provided by an exemplary embodiment of the present application. The device includes:
[0165] The main operation module 901 is used to divide the sub-operations according to the number of sub-operations of the seismic data in the main operation, taking the seismic gather as the minimum division unit, and obtain the sub-operation information of each sub-operation. The seismic data is stored in a shared disk, and the main operation module 901 and each slave operation module 902 support access to the shared disk;
[0166] The master operation module 901 is used to allocate the sub-operations to the slave operation modules 902 according to the node processing capabilities of the slave operation modules 902;
[0167] The sub-job module 902 is used to extract sub-job seismic data from the seismic data according to the sub-job information of the sub-job, and perform data processing on the sub-job seismic data to obtain a sub-job processing result;
[0168] The sub-job module 902 is used to write the sub-job processing result into the shared disk;
[0169] The main operation module 901 is used to merge the sub-operation processing results of each sub-operation when all sub-operations are processed, so as to obtain the main operation data processing result of the seismic data.
[0170] Optionally, the main operation module 901 is further used to:
[0171] Determine the total number of seismic trace data according to the index keyword of the seismic data and the range of the main operation index keyword;
[0172] The main operation module 901 is used to determine the number of reference channels of seismic channel data of the sub-operation according to the total number of seismic channel data and the number of sub-operation divisions;
[0173] The main operation module 901 is used to perform sub-operation division based on the number of seismic data reference channels and the number of seismic data channels contained in each seismic channel gather, taking the seismic channel gather as the minimum division unit, and obtain the sub-operation information of each sub-operation.
[0174] Optionally, the main operation module 901 is further used to:
[0175] In the case where the sum of the number of first seismic channel data from the i-th seismic channel gather to the i+j-th seismic channel gather is less than the number of reference seismic channel data, the sum of the number of second seismic channel data from the i-th seismic channel gather to the i+j+1-th seismic channel gather is greater than the number of reference seismic channel data, and the difference between the sum of the second seismic channel data and the number of reference seismic channel data is greater than a threshold, determining a sub-job index keyword range in the sub-job information according to the i-th seismic channel gather to the i+j-th seismic channel gather;
[0176] The main operation module 901 is used to determine the sub-operation index keyword range in the sub-operation information according to the i-th seismic track gather to the i+j+1-th seismic track gather, when the sum of the first seismic track data channel numbers from the i-th seismic track gather to the i+j+1-th seismic track gather is less than the seismic track data reference channel number, the sum of the second seismic track data channel numbers from the i-th seismic track gather to the i+j+1-th seismic track gather is greater than the seismic track data reference channel number, and the difference between the sum of the second seismic track data channel numbers and the seismic track data reference channel number is less than a threshold value;
[0177] The slave operation module 902 is further used for:
[0178] The sub-job seismic data is extracted from the seismic data according to the index keyword and the sub-job index keyword range contained in the sub-job information.
[0179] Optionally, the shared disk stores a sub-job status file, wherein the sub-job status file contains the running status of each sub-job, and the running status includes a waiting state, a running state, a running success state, a running failure state, and a canceled running state;
[0180] The main operation module 901 is also used for:
[0181] According to the node processing capability of each slave job module 902, the sub-job in a waiting-to-run state is allocated to the slave job module 902;
[0182] The slave operation module 902 is further used for:
[0183] The running state of the sub-job is switched from the waiting state to the running state.
[0184] Optionally, the slave operation module 902 is further used to:
[0185] Switching the running state of the sub-job from a running state to a running successful state;
[0186] Optionally, the slave operation module 902 is further used to:
[0187] When the sub-job processing fails, updating the failure count of the sub-job in the sub-job status file;
[0188] The sub-job module 902 is further configured to switch the running state of the sub-job from the running state to the running failure state when the number of failures reaches a number threshold;
[0189] The sub-job module 902 is further configured to switch the running state of the sub-job from the running state to the waiting state when the number of failures does not reach the number threshold.
[0190] Optionally, the main operation module 901 is further used to:
[0191] Obtaining the number of sub-job failures of the sub-jobs in the failed running state from the sub-job status file;
[0192] The main job module 901 is further configured to terminate the running main job and restart the main job when the ratio of the number of failed sub-jobs to the total number of sub-jobs reaches a ratio threshold;
[0193] The slave job node 902 is further used to switch the sub-job in the running state to the canceled running state when the main job is terminated;
[0194] The slave job module 902 is further configured to switch the sub-jobs in the failed operation state and the canceled operation state to the waiting operation state when the master job is restarted.
[0195] Optionally, the main operation module 901 is further used to:
[0196] Obtaining a slave job node work log from the shared disk, wherein the slave job node work log includes a success rate of each slave job module 902 processing a sub-job;
[0197] The master job module 901 is further used to filter the slave job modules 902 with low success rates based on the slave job node work logs, where the success rate of the sub-jobs processed by the slave job modules 902 with low success rates is lower than a success rate threshold.
[0198] Optionally, the main operation module 901 is further used to:
[0199] Obtaining a main job identifier of the main job;
[0200] The main job module 901 is further configured to create a working directory for the main job when the working directory corresponding to the main job identifier does not exist in the shared disk, and the working directory is used to store files related to the main job and the sub-jobs;
[0201] The main operation module 901 is further configured to continue using the working directory when a working directory corresponding to the main operation identifier exists in the shared disk.
[0202] It should be noted that: the device provided in the above embodiment is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.
[0203] Please refer to Fig.10 , which shows a computer system structure diagram provided by an exemplary embodiment of the present application. The system includes a master operation node 1010 and a plurality of slave operation nodes 1020.
[0204] The master operation node 1010 is used to divide the sub-operations according to the number of sub-operations of the seismic data in the master operation, taking the seismic gather as the minimum division unit, and obtain the sub-operation information of each sub-operation. The seismic data is stored in a shared disk, and the master operation node 1010 and each of the slave operation nodes 1020 support access to the shared disk;
[0205] The master operation node 1010 is used to allocate the sub-operation to each slave operation node 1020 according to the node processing capability of each slave operation node 1020;
[0206] The slave operation node 1020 is used to extract sub-operation seismic data from the seismic data according to the sub-operation information of the sub-operation, and perform data processing on the sub-operation seismic data to obtain a sub-operation processing result;
[0207] The slave job node 1020 is used to write the sub-job processing result into the shared disk;
[0208] The main operation node 1010 is used to merge the sub-operation processing results of each sub-operation when all sub-operations are processed, so as to obtain the main operation data processing result of the seismic data.
[0209] Optionally, the master operation node 1010 is further used for:
[0210] Determine the total number of seismic trace data according to the index keyword of the seismic data and the range of the main operation index keyword;
[0211] Determining the number of reference channels of seismic channel data of the sub-operation according to the total number of seismic channel data and the number of sub-operation divisions;
[0212] According to the number of seismic data reference channels and the number of seismic data channels contained in each seismic channel gather, sub-operation division is performed with the seismic channel gather as the minimum division unit to obtain the sub-operation information of each sub-operation.
[0213] Optionally, the master operation node 1010 is further used for:
[0214] In the case where the sum of the number of first seismic channel data from the i-th seismic channel gather to the i+j-th seismic channel gather is less than the number of reference seismic channel data, the sum of the number of second seismic channel data from the i-th seismic channel gather to the i+j+1-th seismic channel gather is greater than the number of reference seismic channel data, and the difference between the sum of the second seismic channel data and the number of reference seismic channel data is greater than a threshold, determining a sub-job index keyword range in the sub-job information according to the i-th seismic channel gather to the i+j-th seismic channel gather;
[0215] The main operation node 1010 is further used to determine the sub-operation index keyword range in the sub-operation information according to the i-th seismic track gather to the i+j+1-th seismic track gather, when the sum of the first seismic track data channel numbers of the i-th seismic track gather to the i+j+1-th seismic track gather is less than the seismic track data reference channel number, the sum of the second seismic track data channel numbers of the i-th seismic track gather to the i+j+1-th seismic track gather is greater than the seismic track data reference channel number, and the difference between the sum of the second seismic track data channel numbers and the seismic track data reference channel number is less than a threshold value;
[0216] The slave operation node 1020 is further used for:
[0217] The sub-job seismic data is extracted from the seismic data according to the index keyword and the sub-job index keyword range contained in the sub-job information.
[0218] Optionally, the master operation node 1010 is further used for:
[0219] Allocate the sub-job in a waiting-to-run state to the slave job node 1020 according to the node processing capacity of each slave job node 1020;
[0220] The slave operation node 1020 is further used for:
[0221] The running state of the sub-job is switched from the waiting state to the running state.
[0222] Optionally, the slave operation node 1020 is further used for:
[0223] Switching the running state of the sub-job from a running state to a running successful state;
[0224] The slave job node 1020 is further configured to update the failure count of the sub-job in the sub-job status file when the sub-job fails to be processed;
[0225] The slave job node 1020 is further configured to switch the running state of the sub-job from the running state to the running failed state when the number of failures reaches a number threshold;
[0226] The slave job node 1020 is further configured to switch the running state of the sub-job from the running state to the waiting-to-run state when the number of failures does not reach the number threshold.
[0227] Optionally, the master operation node 1010 is further used for:
[0228] Obtaining the number of sub-job failures of the sub-jobs in the failed running state from the sub-job status file;
[0229] The main job node 1010 is further configured to terminate the running main job and restart the main job when the ratio of the number of failed sub-jobs to the total number of sub-jobs reaches a ratio threshold;
[0230] The slave job node 1020 is further used to switch the sub-job in the running state to the canceled running state when the main job stops running;
[0231] The slave job node 1020 is further configured to switch the sub-jobs in the failed operation state and the canceled operation state to the waiting state when the master job is restarted.
[0232] Optionally, the master operation node 1010 is further used for:
[0233] Obtaining a slave job node work log from the shared disk, wherein the slave job node work log includes a success rate of each slave job node 1020 processing a sub-job;
[0234] The master operation node 1010 is further used to filter the slave operation node 1020 with low success rate based on the slave operation node work log, where the success rate of processing sub-operations of the slave operation node 1020 with low success rate is lower than the success rate threshold.
[0235] Optionally, the master operation node 1010 is further used for:
[0236] Obtaining a main job identifier of the main job;
[0237] The main job node 1010 is further configured to create a working directory for the main job when the working directory corresponding to the main job identifier does not exist in the shared disk, and the working directory is used to store files related to the main job and the sub-jobs;
[0238] The master operation node 1010 is further configured to continue using the working directory when a working directory corresponding to the master operation identifier exists in the shared disk.
[0239] It should be noted that the computer system embodiment and the method embodiment provided in the above embodiments belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.
[0240] Please refer to Fig.11 , which shows a structural block diagram of a computer device provided by an exemplary embodiment of the present application. The computer device includes one or more of the following components: a processor 1110 and a memory 1120.
[0241] Optionally, the processor 1110 utilizes various interfaces and lines to connect various parts within the entire computer device, and performs various functions of the computer device and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory 1120, and calling data stored in the memory 1120.
[0242] The memory 1120 may include a random access memory (RAM) or a read-only memory (ROM). The memory 1120 may be used to store instructions, programs, codes, code sets or instruction sets.
[0243] An embodiment of the present application further provides a computer-readable storage medium, which stores at least one program code, and the at least one program code is used to be executed by a processor to implement the seismic operation execution method as described in the above embodiment.
[0244] The above description is only an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A method for performing seismic operations, It is characterized in that The method comprises: The master operation node divides the sub-operation into sub-operations according to the number of sub-operation divisions of the seismic data in the master operation, taking the seismic gather as the minimum division unit, and obtains the sub-operation information of each sub-operation, wherein the seismic data is stored in a shared disk, and the master operation node and each slave operation node support access to the shared disk; The master operation node allocates the sub-operation to the slave operation node according to the node processing capability of each slave operation node; The slave operation node extracts sub-operation seismic data from the seismic data according to the sub-operation information of the sub-operation, and performs data processing on the sub-operation seismic data to obtain a sub-operation processing result; The slave job node writes the sub-job processing result into the shared disk; When all sub-jobs are processed, the main operation node merges the sub-operation processing results of each sub-job to obtain the main operation data processing result of the seismic data.
2. The method according to claim 1, It is characterized in that The main operation node divides the sub-operations according to the number of sub-operations of the seismic data in the main operation, taking the seismic gather as the minimum division unit, and obtains the sub-operation information of each sub-operation, including: The main operation node determines the total number of seismic trace data according to the index keyword of the seismic data and the main operation index keyword range; The main operation node determines the number of reference channels of the seismic channel data of the sub-operation according to the total number of the seismic channel data and the number of sub-operation divisions; The main operation node divides the sub-operations according to the number of seismic channel data reference channels and the number of seismic channel data channels contained in each seismic channel gather, taking the seismic channel gather as the minimum division unit, and obtains the sub-operation information of each sub-operation.
3. The method according to claim 2, It is characterized in that The main operation node divides the sub-operation into sub-operations according to the number of seismic channel data reference channels and the number of seismic channel data channels included in each seismic channel gather, taking the seismic channel gather as the minimum division unit, and obtains the sub-operation information of each sub-operation, including: In the case where the sum of the number of first seismic channel data from the i-th seismic channel gather to the i+j-th seismic channel gather is less than the number of reference seismic channel data, the sum of the number of second seismic channel data from the i-th seismic channel gather to the i+j+1-th seismic channel gather is greater than the number of reference seismic channel data, and the difference between the sum of the second seismic channel data and the number of reference seismic channel data is greater than a threshold, the main operation node determines the sub-operation index keyword range in the sub-operation information according to the i-th seismic channel gather to the i+j-th seismic channel gather; In the case where the sum of the number of first seismic channel data from the i-th seismic channel gather to the i+j-th seismic channel gather is less than the number of reference seismic channel data, the sum of the number of second seismic channel data from the i-th seismic channel gather to the i+j+1-th seismic channel gather is greater than the number of reference seismic channel data, and the difference between the sum of the second seismic channel data and the number of reference seismic channel data is less than a threshold, the main operation node determines the sub-operation index keyword range in the sub-operation information according to the i-th seismic channel gather to the i+j+1-th seismic channel gather; The slave operation node extracts sub-operation seismic data from the seismic data according to the sub-operation information of the sub-operation, including: The slave operation node extracts the sub-operation seismic data from the seismic data according to the index keyword and the sub-operation index keyword range contained in the sub-operation information.
4. The method according to claim 1, It is characterized in that The shared disk stores a sub-job status file, which contains the running status of each sub-job, including a waiting state, a running state, a successful running state, a failed running state, and a canceled running state; The master operation node allocates the sub-operation to the slave operation node according to the node processing capability of each slave operation node, including: The master operation node allocates the sub-operation in a waiting-to-run state to the slave operation node according to the node processing capacity of each slave operation node; After the master operation node allocates the sub-operation to the slave operation node according to the node processing capability of each slave operation node, the method includes: The slave job node switches the running state of the sub-job from a waiting-to-run state to a running state.
5. The method according to claim 4, It is characterized in that After the slave job node writes the sub-job processing result into the shared disk, the method further includes: The slave job node switches the running state of the sub-job from a running state to a running success state; The method further comprises: In the case where the sub-job processing fails, the slave job node updates the failure count of the sub-job in the sub-job status file; When the number of failures reaches a threshold, the slave job node switches the running state of the sub-job from a running state to a running failure state; When the number of failures does not reach the number threshold, the slave job node switches the running state of the sub-job from the running state to the waiting state.
6. The method according to claim 5, It is characterized in that The method further comprises: The main job node obtains the number of sub-job failures of the sub-jobs in the failed running state from the sub-job status file; When the proportion of the number of sub-job failures to the total number of sub-jobs reaches a proportion threshold, the main job node terminates the running main job and restarts the main job; When the main job stops running, the slave job node switches the sub-job in the running state to the canceled running state; When the main job is restarted, the slave job node switches the sub-jobs in the failed operation state and the canceled operation state to the waiting state.
7. The method according to claim 6, It is characterized in that Before the main operation node restarts the main operation, the method further includes: The master operation node obtains the slave operation node work log from the shared disk, wherein the slave operation node work log includes the success rate of each slave operation node in processing the sub-operation; The master operation node filters the slave operation nodes with low success rates based on the work logs of the slave operation nodes, where the success rate of processing sub-operations of the slave operation nodes with low success rates is lower than a success rate threshold.
8. The method according to claim 1, It is characterized in that The method further comprises: The main operation node obtains the main operation identifier of the main operation; In the case that the working directory corresponding to the main job identifier does not exist in the shared disk, the main job node creates a working directory for the main job, and the working directory is used to store files related to the main job and the sub-jobs; In the case that the working directory corresponding to the main job identifier exists in the shared disk, the main job node continues to use the working directory.
9. A seismic operation execution device, It is characterized in that The device comprises: A main operation module, used to divide the sub-operation into sub-operations according to the number of sub-operation divisions of the seismic data in the main operation, taking the seismic gather as the minimum division unit, and obtain the sub-operation information of each sub-operation, wherein the seismic data is stored in a shared disk, and the main operation module and each slave operation module support access to the shared disk; The master operation module is used to allocate the sub-operations to the slave operation modules according to the node processing capabilities of the slave operation modules; The sub-job module is used to extract sub-job seismic data from the seismic data according to the sub-job information of the sub-job, and perform data processing on the sub-job seismic data to obtain a sub-job processing result; The slave job module is used to write the sub-job processing result into the shared disk; The main operation module is used to merge the sub-operation processing results of each sub-operation when all sub-operations are processed, so as to obtain the main operation data processing result of the seismic data.
10. A computer system, It is characterized in that The computer system includes a master operation node and a plurality of slave operation nodes; The master operation node is used to divide the sub-operations according to the number of sub-operations of the seismic data in the master operation, taking the seismic gather as the minimum division unit, and obtain the sub-operation information of each sub-operation. The seismic data is stored in a shared disk, and the master operation node and each of the slave operation nodes support access to the shared disk; The master operation node is used to allocate the sub-operations to the slave operation nodes according to the node processing capabilities of the slave operation nodes; The slave operation node is used to extract sub-operation seismic data from the seismic data according to the sub-operation information of the sub-operation, and perform data processing on the sub-operation seismic data to obtain a sub-operation processing result; The slave job node is used to write the sub-job processing result into the shared disk; The main operation node is used to merge the sub-operation processing results of each sub-operation when all sub-operations are processed, so as to obtain the main operation data processing result of the seismic data.
Citation Information
Cited By
Seismic data preprocessing method and device, storage medium and electronic equipment
CN120908866A