Drilling radar data distributed processing system and method based on Apache Spark

Through the distributed processing system of drilling radar data based on Apache Spark, the problem that traditional methods are difficult to deal with large-scale drilling radar data is solved, efficient parallel processing and in-depth analysis are achieved, and real-time data flow and powerful machine learning capabilities are supported.

CN120405775APending Publication Date: 2025-08-01XIAN RES INST OF CHINA COAL TECH & ENG GRP CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510354764.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

Traditional drilling radar data processing methods are difficult to cope with the real-time processing requirements of large-scale drilling radar data sets, which are limited by computing power and data processing efficiency.

Method used

Apache Spark-based drilling radar data distributed processing system, including data collection, synchronization, storage, computing, interface, application and coordination modules, uses Spark Streaming, MLlib, Graphx, Spark Core and other components for data preprocessing, feature extraction and machine learning inversion analysis.

Benefits of technology

It realizes parallel processing of large-scale drilling radar data, significantly improving data processing efficiency and analysis depth, supporting real-time data stream processing and powerful machine learning capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120405775A_ABST
    Figure CN120405775A_ABST
Patent Text Reader

Abstract

The invention relates to a drilling radar data distributed processing system and method based on Apache Spark. The system comprises a data collection module, a data synchronization module, a data storage module, a data calculation module, a data warehouse, a data interface module, a data application module, a uniform resource scheduling management module and a distributed coordination module. According to the method, large-scale drilling radar data parallel processing can be realized, and the data processing efficiency and the analysis depth are remarkably improved by optimizing the data processing flow.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of borehole radar data processing. Specifically, it relates to a distributed processing system and method for borehole radar data based on Apache Spark. Background Art

[0002] With the application of borehole radar detection technology in the construction of transparent working faces in coal mines, the borehole radar instrumentation has become increasingly mature, the applied borehole depth has become deeper, the number of boreholes has increased, and the borehole radar data has shown a blowout growth, gradually entering the era of big data for borehole radar. Traditional methods for processing borehole radar data are often limited by computing power and data processing efficiency, and it is difficult to meet the real-time processing requirements of large-scale borehole radar data sets. Summary of the Invention

[0003] In order to overcome at least one deficiency in the prior art, this application provides a distributed processing system and method for borehole radar data based on Apache Spark.

[0004] In a first aspect, a distributed processing system for borehole radar data based on Apache Spark is provided, including: a data collection module, a data synchronization module, a data storage module, a data calculation module, a data warehouse, a data interface module, a data application module, a unified resource scheduling and management module, and a distributed coordination module;

[0005] The data collection module is used to collect raw acquisition data, where the raw acquisition data includes data collected by borehole radar instruments, geological environment data of the exploration work area, and other geophysical exploration data of the exploration work area;

[0006] The data synchronization module is used to achieve synchronous transmission of the raw acquisition data;

[0007] The data storage module is used to store relevant data during the distributed processing of borehole radar data;

[0008] The data calculation module is used to process the raw acquisition data to determine the abnormal distribution detected by the borehole radar;

[0009] The data warehouse is used to store the raw acquisition data and the data during the data processing by the data calculation module;

[0010] The data interface module is used to design transmission protocols, interface specifications, and data formats to achieve encryption and authentication during data transmission;

[0011] The data application module is used to achieve front-end display, including report statistics on the data stored in the data warehouse, display on a large-screen cockpit, data management, and production interpretation reports;

[0012] The unified resource scheduling and management module is used for the management of cluster resources in a distributed processing system and the scheduling of jobs;

[0013] The distributed coordination module is used to maintain the configuration information of the distributed processing system and provide distributed synchronization and group services for data processing.

[0014] In one embodiment, the data storage module includes the message queue Kafka, the relational database Postgres, the in-memory database Redis, the distributed data file system HDFS, and the data warehouse tool Hive;

[0015] The message queue Kafka is used for asynchronous data transmission between different components in the distributed processing system. The relational database Postgres is used to store structured data. The in-memory database Redis stores data in RAM to achieve fast access. The distributed data file system HDFS is used to store large-scale data sets. The data warehouse tool Hive is built on the Hadoop ecosystem, uses the distributed data file system HDFS as the underlying storage, and allows the front end of the distributed processing system to operate and analyze the data stored on the distributed data file system HDFS through the HiveQL query language.

[0016] In one embodiment, the data calculation module includes Spark Streaming, MLlib, Graphx, and SparkCore. Spark Streaming is used to preprocess the original collected data to obtain preprocessed data; MLlib is used to extract features from the preprocessed data and perform inversion analysis based on machine learning to determine the abnormal distribution detected by the borehole radar; Graphx is used for parallel processing of the images formed by the borehole radar data; Spark Core is used for parallel processing of large-scale data sets, optimizing resource allocation and task scheduling.

[0017] In one embodiment, the data processing of the original collected data to determine the abnormal distribution detected by the borehole radar includes:

[0018] Preprocess the original collected data to obtain preprocessed data; the preprocessing includes median filtering, band-pass filtering, wavelet threshold denoising, automatic gain processing, and normalization processing;

[0019] Extract features from the preprocessed data to obtain time-domain amplitude features, frequency-domain features, spatial features, and deep learning automatic features; the time-domain amplitude features include peak amplitude, average amplitude, integral of the reflected wave envelope energy, and direct wave arrival time; the frequency-domain features include main frequency, frequency band energy ratio, and multi-scale energy distribution; the spatial features include profile continuity and amplitude gradient along the borehole direction; the deep learning automatic features include the features extracted by the convolutional layer;

[0020] Input the preprocessed data and the extracted features into the machine learning model to inversely obtain the abnormal distribution detected by the borehole radar.

[0021] In a second aspect, a distributed processing method for borehole radar data based on Apache Spark is provided, including:

[0022] Step 1, the user uploads the original collected data, and after being processed by the data synchronization module, it is uploaded to the data storage module; the original collected data includes the data collected by the borehole radar instrument, the geological environment data of the detection work area, and other geophysical exploration data of the detection work area;

[0023] Step 2, the unified resource scheduling and management module manages the data calculation module, and the data calculation module processes the original collected data to determine the abnormal distribution detected by the borehole radar;

[0024] Step 3, the data processed by the data calculation module and the original collected data uploaded by the user are transmitted to the data warehouse together, and the data warehouse summarizes according to the dimension data and the detailed data, and converts the data into application data;

[0025] Step 4, the data application module outputs the data stored in the data warehouse according to report statistics, interpretation reports, data management, and large-screen cockpits respectively, and transmits them to the user for the user to view.

[0026] In one embodiment, the data calculation module processes the original collected data to determine the abnormal distribution detected by the borehole radar, including:

[0027] Preprocess the original collected data to obtain the preprocessed data; the preprocessing includes median filtering, band-pass filtering, wavelet threshold denoising, automatic gain processing, and normalization processing;

[0028] Extract features from the preprocessed data to obtain time-domain amplitude features, frequency-domain features, spatial features, and deep learning automatic features; the time-domain amplitude features include peak amplitude, average amplitude, integral of reflected wave envelope energy, and direct wave arrival time; the frequency-domain features include main frequency, frequency band energy ratio, and multi-scale energy distribution; the spatial features include profile continuity and amplitude gradient along the borehole direction; the deep learning automatic features include the features extracted by the convolutional layer;

[0029] Input the preprocessed data and the extracted features into the machine learning model to inversely obtain the abnormal distribution detected by the borehole radar.

[0030] Compared with the prior art, the present application has the following beneficial effects: The distributed processing system and method for borehole radar data based on Apache Spark of the present application can achieve parallel processing of large-scale borehole radar data, and significantly improve the efficiency of data processing and the depth of analysis by optimizing the data processing flow. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The present application can be better understood by referring to the following description in conjunction with the accompanying drawings. The drawings, together with the following detailed description, are included in this specification and form a part of this specification. In the drawings:

[0032] Figure 1 The structural block diagram of the distributed processing system for borehole radar data based on Apache Spark is shown;

[0033] Figure 2 The flow block diagram of the distributed processing method for borehole radar data based on Apache Spark is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0034] Hereinafter, exemplary embodiments of the present application will be described in conjunction with the accompanying drawings. For the sake of clarity and conciseness, not all features of the actual embodiments are described in the specification. However, it should be understood that many specific decisions specific to the embodiments may be made during the development of any such actual embodiment in order to achieve the specific goals of the developer, and these decisions may vary with different embodiments.

[0035] Here, it should also be noted that in order to avoid obscuring the present application with unnecessary details, only the device structures closely related to the solution of the present application are shown in the drawings, and other details less related to the present application are omitted.

[0036] It should be understood that the present application is not limited to the described embodiments only due to the following description with reference to the drawings. In this document, where feasible, the embodiments can be combined with each other, features can be replaced or borrowed between different embodiments, and one or more features can be omitted in one embodiment.

[0037] Apache Spark is a distributed in-memory computing framework that can load data into memory for processing. Compared with traditional disk-based processing models (such as Hadoop MapReduce), it greatly improves the processing speed, especially when dealing with iterative computing tasks (such as machine learning training). To solve the problem of rapid processing of large-scale borehole radar datasets, this application provides a distributed processing system and method for borehole radar data based on Apache Spark, which can achieve parallel processing of large-scale borehole radar data and significantly improve the efficiency of data processing and the depth of analysis by optimizing the data processing flow.

[0038] An embodiment of this application provides a distributed processing system for borehole radar data based on Apache Spark. Figure 1 The structural block diagram of the distributed processing system for borehole radar data based on Apache Spark is shown. Refer to Figure 1 , the system includes: a data collection module, a data synchronization module, a data storage module, a data calculation module, a data warehouse, a data interface module, a data application module, a unified resource scheduling and management module, and a distributed coordination module; the specific implementation functions of each module are introduced in detail below.

[0039] The data collection module is used to collect raw acquisition data, and the raw acquisition data includes data collected by borehole radar instruments, geological environment data of the exploration area, and other geophysical exploration data of the exploration area.

[0040] Among them, the collection of data collected by borehole radar instruments includes: selecting data collected by high-precision borehole radar equipment to ensure the accuracy and integrity of data collection; designing a data collection interface to achieve real-time data transmission between the equipment and the data processing system; developing data collection software, including data formatting, timestamp recording, and quality control.

[0041] The geological environment data of the exploration area includes the formation lithology situation where the borehole radar borehole is located, the geological tectonic movement situation of the exploration area, and the mining map of the exploration area.

[0042] Other geophysical exploration data of the exploration area includes data from surface or underground exploration such as seismic exploration and electrical prospecting.

[0043] The data synchronization module is used to achieve the synchronous transmission of the original collected data. It mainly includes: configuring the DataX or Sqoop tool to achieve data synchronization between HDFS and other data storage systems. Developing a custom data transmission service based on Spark, which supports multiple data formats and protocols, including the borehole radar data format, the text of the geological data in the work area, cad, picture data information format, and the data formats of other exploration information in the work area, such as the segy data format of seismic exploration and the dat format of electrical prospecting, to achieve the monitoring and logging of the data transmission of the borehole radar instrument collected data and the data collected in the work area, and ensure the consistency and integrity of the data.

[0044] The data storage module is used to store the relevant data during the distributed processing of the borehole radar data.

[0045] Specifically, the data storage module includes the message queue Kafka, the relational database Postgres, the in-memory database Redis, the distributed data file system HDFS, and the data warehouse tool Hive.

[0046] The message queue Kafka is used for asynchronous data transmission between different components in the distributed processing system, including the response message queue of the front-end users and the message queue for data transmission between other modules. In the data pipeline, the borehole radar data flows into Hive or other modules for storage and analysis.

[0047] The relational database Postgres is used to store structured data, mainly including the original collected data of the borehole radar, the data after each processing step, and the structured data such as the processing operation commands initiated by the front end. It supports SQL queries and can be integrated with Hive. For example, data can be exported from Hive to the relational database or imported from the relational database to Hive through the Sqoop tool.

[0048] The in-memory database Redis stores data in the RAM to achieve fast access. It mainly stores data in the PC hardware where the front end is located, and is used to cache the results of Hive queries or as a fast storage solution for real-time data processing.

[0049] The distributed data file system HDFS is used to store large-scale data sets. The borehole radar-related data, the geological information of the work area, and other exploration materials in the work area are stored in different data set forms, and real-time data backup is performed on the entire relevant data of the distributed processing system. The backup data is stored in the form of a data set for users to recover at any time in case of special situations, so as to ensure the security and reliability of the data of the borehole radar intelligent analysis system.

[0050] The data warehouse tool Hive is built on top of the Hadoop ecosystem. It uses the distributed data file system HDFS as the underlying storage, allowing the front-end of the distributed processing system to operate on and analyze the data stored on the distributed data file system HDFS through the HiveQL query language.

[0051] The data calculation module is used to process the raw collected data and determine the abnormal distribution detected by the borehole radar.

[0052] Specifically, the data calculation module includes Spark Streaming, MLlib, Graphx, and Spark Core.

[0053] Spark Streaming is used to preprocess the raw collected data to obtain the preprocessed data; and to perform preliminary processing on the real-time data stream uploaded by the front-end, including filtering, denoising, and data normalization.

[0054] MLlib is used to extract features from the preprocessed data and perform inversion analysis based on machine learning to determine the abnormal distribution detected by the borehole radar; MLlib is a module specifically for machine learning processing of borehole radar data, and various machine learning methods can be designed according to needs, including methods for classifying borehole radar data according to different regions, clustering analysis of data in the same region according to lithology and borehole characteristics, and inversion of preprocessed borehole radar data using machine learning, etc.

[0055] Graphx is used to perform parallel processing on the images formed by borehole radar data; Graphx is a component in Apache Spark. Using this component is mainly to process the images formed by borehole radar data, and use graph algorithms to process the images formed by borehole radar processed data. It can perform large-scale parallel processing on all borehole radar data images in data processing, including enhancement processing of weak information in borehole radar images, edge detection of borehole radar images to intelligently identify structural anomalies, and methods for classifying and learning borehole radar image anomalies, etc.

[0056] Spark Core is used for parallel processing of large-scale data sets, optimizing resource allocation and task scheduling.

[0057] The data warehouse is used to store the originally collected data and the data during the data processing by the data calculation module. Here, the data warehouse is a subsystem for reporting and data analysis in the borehole radar distributed processing system. This data warehouse stores all the original data uploaded by the front-end in the distributed processing system, the data during the data processing, the borehole radar feature extraction data, the borehole radar inversion analysis data, the application data for front-end display, etc. Among them, the original data includes the data measured by the borehole radar instrument, the geological data of the work area, and all the data of other geophysical exploration results in the work area.

[0058] The data interface module is used to design the transmission protocol, interface specification, and data format to achieve encryption and authentication during the data transmission process. The data interface mainly designs the transmission protocol Http / TCP, the interface specification Restful API, and the data format JSON to ensure the encryption and authentication methods during the data transmission process. The transmission protocol Http / TCP is used to transmit data between the front-end client and the server. The interface specification Restful API is the style of designing the borehole radar intelligent analysis system, using standard HTTP methods to perform operations on resources. The data format JSON is the format of data transmission between the front-end and the server.

[0059] The data application module is used to implement front-end display, including report statistics on the data stored in the data warehouse, display on the large-screen cockpit, data management, and production interpretation report. Among them, the report statistics are carried out according to the region, the borehole radar instrument model, and the lithology of the exploration borehole. The display on the large-screen cockpit mainly shows several types of data such as the original data, processed data, inverted data, and interpreted result images detected by the borehole radar on the front-end display. Data management is the management of data by the client, including the upload and deletion of borehole radar data, the management of customer information data, etc. The interpretation report is directly generated by the front-end according to the customer's needs after the borehole radar undergoes a series of data processing and interpretation.

[0060] The unified resource scheduling and management module is used for the management of the cluster resources of the distributed processing system and the scheduling of jobs.

[0061] The distributed coordination Zookeeper module is used to maintain the configuration information of the distributed processing system and provide distributed synchronization and group services for data processing. Apache ZooKeeper is used for cluster coordination to ensure the high availability of services. YARN is configured as the resource manager to optimize the allocation and scheduling of computing resources. A monitoring and logging system, such as Ganglia, Nagios, or a custom solution, is implemented to monitor the cluster performance and status.

[0062] Here, the unified resource scheduling and management Yarn and the distributed coordination Zookeeper work together to manage the operation of the entire borehole radar distributed processing system.

[0063] In one embodiment, data processing is performed on the original collected data to determine the abnormal distribution detected by the borehole radar, including:

[0064] First, preprocess the original collected data to obtain preprocessed data; the preprocessing includes median filtering, band-pass filtering, wavelet threshold denoising, automatic gain processing, and normalization processing; here, a median filter is used for median filtering, a band-pass filter is used for band-pass filtering, a wavelet denoiser is used for wavelet threshold denoising, an automatic gain controller is used for automatic gain processing, and a maximum-minimum normalization processor is used for normalization processing.

[0065] Then, extract features from the preprocessed data to obtain time-domain amplitude features, frequency-domain features, spatial features, and deep learning automatic features; the time-domain amplitude features include peak amplitude, average amplitude, integral of reflected wave envelope energy, and direct wave arrival time; the frequency-domain features include main frequency, frequency band energy ratio, and multi-scale energy distribution; the spatial features include profile continuity and amplitude gradient along the borehole direction; the deep learning automatic features include features extracted by the convolutional layer;

[0066] Here, a maximum value extractor, an average value extractor, an envelope signal energy integral extractor, and a direct wave arrival time automatic extractor are respectively used to extract the peak amplitude, average amplitude, integral of reflected wave envelope energy, and direct wave arrival time;

[0067] Here, the frequency and amplitude information in the frequency domain after Fourier transform is obtained for the processed data, and then a maximum value extractor after Fourier transform is used to extract the main frequency; a frequency band energy ratio extractor is used to extract the frequency band energy ratio; a wavelet transform multi-scale energy extractor is used for the preprocessed data to extract the multi-scale energy distribution;

[0068] A similarity extractor is used for the preprocessed data to extract the profile continuity; an amplitude gradient extractor along the borehole direction is used, and the amplitude gradient is calculated by the first-order difference method to extract the amplitude gradient along the borehole direction; a convolutional layer feature extractor is used to extract the features of the convolutional layer, including local texture and structural features.

[0069] Then, the preprocessed data and the extracted various features are input into the machine learning model to inversely obtain the abnormal distribution detected by the borehole radar.

[0070] Here, the preprocessed data and the extracted features are used as the input of machine learning, and the measured formation dielectric constant is used as the output of machine learning for training to obtain a trained machine learning model; the dielectric constant distribution information of the formation is inverted according to the trained machine learning model, that is, the abnormal distribution detected by the borehole radar. Specifically, supervised learning regression methods such as random forest, support vector machine SVR, deep learning neural network DNN, etc. are used for the machine learning method.

[0071] The embodiment of the present application also provides a distributed processing method for borehole radar data based on Apache Spark. Figure 2 The flowchart of the distributed processing method for borehole radar data based on Apache Spark is shown. Refer to Figure 2 , which mainly includes the following steps:

[0072] Step 1, the user uploads the original collected data, and after being processed by the data synchronization module, it is uploaded to the data storage module; the original collected data includes the data collected by the borehole radar instrument, the geological environment data of the detection work area, and other geophysical exploration data of the detection work area.

[0073] Here, the user uploads the original collected data, which is processed by the data synchronization components DataX\Sqoop\Kettle in the data synchronization module, and then transmitted to the distributed file system HDFS.

[0074] Step 2, the unified resource scheduling and management module manages the data calculation module, and the data calculation module processes the original collected data to determine the abnormal distribution detected by the borehole radar.

[0075] Here, the unified resource scheduling and management module Yarn manages the Spark data, and the Spark data processing processes the data collected by the borehole radar instrument in the distributed file system HDFS in combination with the geological environment data of the work area and other geophysical exploration data of the work area. First, data preprocessing is performed, Spark Streaming is used to process real-time data streams, denoising and data correction are performed, and MLlib is used to extract features and perform inversion analysis on the borehole radar data based on machine learning.

[0076] Step 3, the data processed by the data calculation module and the original collected data uploaded by the user are transmitted to the data warehouse together, and the data warehouse summarizes according to the dimension data and detailed data and converts the data into application data;

[0077] Step 4, the data application module outputs the data stored in the data warehouse according to report statistics, interpretation reports, data management, and large-screen cockpits respectively, and transmits it to the user for the user to view.

[0078] In summary, the present application has the following technical effects:

[0079] 1. By leveraging the high-performance computing power of Spark, the processing speed of borehole radar data has been significantly improved.

[0080] 2. The real-time data stream processing ability enables the system to quickly respond to geological changes.

[0081] 3. The powerful machine learning ability has enhanced the accuracy of intelligent data analysis and pattern recognition.

[0082] 4. Flexible data interfaces and visualization tools have enhanced the accessibility and intuitiveness of data.

[0083] The above are only various implementation manners of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily conceive of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A distributed processing system for borehole radar data based on Apache Spark, characterized in that Including: A data collection module, a data synchronization module, a data storage module, a data calculation module, a data warehouse, a data interface module, a data application module, a unified resource scheduling and management module, and a distributed coordination module; The data collection module is used to collect original acquisition data, and the original acquisition data includes data collected by a borehole radar instrument, geological environment data of the exploration work area, and other geophysical exploration data of the exploration work area; The data synchronization module is used to realize the synchronous transmission of the original acquisition data; The data storage module is used to store relevant data during the distributed processing process of borehole radar data; The data calculation module is used to process the original acquisition data to determine the abnormal distribution detected by the borehole radar; The data warehouse is used to store the original acquisition data and the data during the data processing process of the data calculation module; The data interface module is used to design a transmission protocol, interface specifications, and data formats to realize encryption and authentication during the data transmission process; The data application module is used to realize front-end display, including report statistics on the data stored in the data warehouse, display of a large-screen cockpit, data management, and production interpretation reports; The unified resource scheduling and management module is used to manage the cluster resources of the distributed processing system and schedule jobs; The distributed coordination module is used to maintain the configuration information of the distributed processing system and provide distributed synchronization and group services for data processing.

2. The system according to claim 1, wherein The data storage module includes a message queue Kafka, a relational database Postgres, a memory database Redis, a distributed data file system HDFS, and a data warehouse tool Hive; The message queue Kafka is used for asynchronous data transmission between different components in the distributed processing system. The relational database Postgres is used to store structured data. The memory database Redis stores data in RAM to achieve fast access. The distributed data file system HDFS is used to store large-scale data sets. The data warehouse tool Hive is built on the Hadoop ecosystem, uses the distributed data file system HDFS as the underlying storage, and allows the front-end of the distributed processing system to operate and analyze the data stored on the distributed data file system HDFS through the HiveQL query language.

3. The system according to claim 1, characterized in that, The data calculation module includes Spark Streaming, MLlib, Graphx, and Spark Core. Spark Streaming is used to preprocess the original acquisition data to obtain preprocessed data; MLlib is used to extract features from the preprocessed data and perform inversion analysis based on machine learning to determine the abnormal distribution detected by the borehole radar. Graphx is used for parallel processing of the images formed by the borehole radar data. Spark Core is used for parallel processing of large-scale data sets to optimize resource allocation and task scheduling.

4. The system according to claim 1, wherein The processing of the original acquisition data to determine the abnormal distribution detected by the borehole radar includes: Preprocess the original acquired data to obtain the preprocessed data; the preprocessing includes median filtering, band-pass filtering, wavelet threshold denoising, automatic gain processing, and normalization processing; Extract features from the preprocessed data to obtain time-domain amplitude features, frequency-domain features, spatial features, and deep learning automatic features; the time-domain amplitude features include peak amplitude, average amplitude, integral of reflected wave envelope energy, and direct wave arrival time; the frequency-domain features include main frequency, frequency band energy ratio, and multi-scale energy distribution; the spatial features include profile continuity and amplitude gradient along the borehole direction; the deep learning automatic features include features extracted by the convolutional layer; Input the preprocessed data and the extracted features into a machine learning model to inversely obtain the abnormal distribution detected by the borehole radar.

5. A distributed processing method for borehole radar data based on Apache Spark, characterized in that, It includes: Step 1: The user uploads the original acquired data. After being processed by the data synchronization module, it is uploaded to the data storage module; the original acquired data includes data collected by the borehole radar instrument, geological environment data of the exploration area, and other geophysical exploration data of the exploration area; Step 2: The unified resource scheduling and management module manages the data calculation module, and the data calculation module processes the original acquired data to determine the abnormal distribution detected by the borehole radar; Step 3: The data processed by the data calculation module and the original acquired data uploaded by the user are transmitted to the data warehouse together. The data warehouse summarizes according to dimension data and detailed data and converts the data into application data; Step 4: The data application module outputs the data stored in the data warehouse according to report statistics, interpretation reports, data management, and large-screen cockpit respectively, and transmits it to the user for the user to view.

6. The method according to claim 5, characterized in that, Among them, The data calculation module processes the original acquired data to determine the abnormal distribution detected by the borehole radar, including: Preprocess the original acquired data to obtain the preprocessed data; the preprocessing includes median filtering, band-pass filtering, wavelet threshold denoising, automatic gain processing, and normalization processing; Extract features from the preprocessed data to obtain time-domain amplitude features, frequency-domain features, spatial features, and deep learning automatic features; the time-domain amplitude features include peak amplitude, average amplitude, integral of reflected wave envelope energy, and direct wave arrival time; the frequency-domain features include main frequency, frequency band energy ratio, and multi-scale energy distribution; the spatial features include profile continuity and amplitude gradient along the borehole direction; the deep learning automatic features include features extracted by the convolutional layer; Input the preprocessed data and the extracted features into a machine learning model to inversely obtain the abnormal distribution detected by the borehole radar.

Citation Information

Patent Citations

  • Method for inverting coal-rock interface by borehole radar based on deep learning

    CN118033544A

  • KR20230133705A