A method and system for real-time prediction of rock properties based on acoustic and vibration signal characteristics during drilling and rock fragmentation

By constructing the AlexNet-RCBAM convolutional neural network, combining the CBAM attention mechanism and residual structure, the acoustic and vibrating signals during drilling into rock crushing are solved, and the subjectivity and hysteresis of traditional lithology recognition methods are achieved, real-time and accuracy of lithology recognition are met, and the lithology recognition needs are met under complex geological conditions.

CN117216676BActive Publication Date: 2025-08-15CHENGDU UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310799393.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-30
Publication Date
2025-08-15
Estimated Expiration
2043-06-30

AI Technical Summary

Technical Problem

In the existing drilling technology, lithology identification methods have strong subjectivity and high lag. lithology identification technology while drilling has high requirements for data acquisition, acquisition amount and timeliness of identification. Traditional methods are difficult to meet the lithology identification needs under complex geological conditions, and the acoustic and vibrating signals are not fully utilized in lithology identification.

Method used

A method is constructed to predict lithologic properties in real-time based on the characteristics of the acoustic and oscillatory signal through the AlexNet-RCBAM convolutional neural network, combined with the CBAM attention mechanism and residual structure, integrate basic models of multiple signal directions, perform data augmentation and training, and build an integrated model, and use pre-emphasis, wavelet threshold noise reduction, bandpass filtering and time-frequency domain analysis to process the acoustic and oscillatory signal.

Benefits of technology

It improves the accuracy and generalization of lithology recognition, realizes real-time and accuracy of lithology recognition, reduces noise interference, enhances the reliability and accuracy of signal processing, and meets the lithology recognition needs under complex geological conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117216676B_ABST
    Figure CN117216676B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for real-time prediction of lithology based on the characteristics of acoustic vibration signals during drilling and rock crushing, belonging to the field of drilling technology. The present invention analyzes the lithology characteristics and drilling status characteristics contained in the acoustic vibration signals, combines them with a deep learning convolutional neural network, constructs a lithology recognition model based on acoustic vibration signals, and explains the decision-making of the model. In view of the different lithology recognition capabilities of each model, combined with the idea of joint decision-making of the models, a multi-model integrated learning intelligent lithology recognition method and system are provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of drilling technology, and in particular relates to a method and system for real-time prediction of lithology based on acoustic vibration signal characteristics during drilling and rock crushing. Background Art

[0002] Geological exploration is one of the foundations of economic and social development and is of great significance to my country's economic development. Drilling, a widely used and reliable exploration technology, can directly obtain geological samples from a region. However, current drilling operations face challenges such as harsh engineering environments, complex geological conditions, and insufficient drilling efficiency. Rapid development of drilling technology is urgently needed to meet engineering needs.

[0003] Intelligent drilling technology is currently a major development direction in drilling technology. It combines artificial intelligence with drilling technology to intelligently optimize drilling parameters, control drilling trajectories, and identify drilling risks, ultimately improving drilling efficiency and reducing construction costs. The development of intelligent drilling technology is inseparable from the research on lithology identification technology. Traditional lithology identification methods in the drilling field mainly rely on core or cutting analysis. Technicians rely on experience supplemented by physical or chemical methods to determine the lithology of the core. This lithology identification method is not only subjective but also has a significant lag. While-drilling lithology identification allows for real-time identification of the lithology at the drill bit during the drilling process, directly providing geological data for exploration operations. Furthermore, based on formation lithology, it is possible to timely select the drill bit type, optimize drilling parameters, and control the drilling trajectory, thereby improving drilling efficiency.

[0004] Research on LWD lithology identification technology is primarily based on measurement while drilling (MWD) and logging while drilling (LWD) data. These data primarily collect drilling parameters such as weight on bit, torque, and rotational speed, as well as rock physical information such as resistivity, natural gamma ray, and neutron density. Using this LWD information to identify formation lithology has achieved some success. However, as the geological conditions facing drilling operations become increasingly complex, LWD lithology identification technology places higher demands on data collectibility, data volume, and timeliness of lithology identification. Conventional LWD data is no longer sufficient.

[0005] Acoustic vibration signals have good propagation properties and rich data features. During drilling, rock is broken by interaction with the drill bit. The acoustic vibration signals generated during this process contain lithologic information. Therefore, research on lithologic identification while drilling based on acoustic vibration signals is undoubtedly of great value. This is of great significance for improving drilling technology and providing a reference and basis for solving engineering problems.

[0006] Although lithology identification methods based on acoustic and vibration signals are beginning to attract research attention, relatively little research has been conducted. Current research primarily analyzes and compares signal characteristics from a single perspective, either acoustic or vibration, to artificially distinguish lithologies. Few lithology identification models have been developed. Even in the few studies that have constructed lithology identification models, their accuracy and generalizability are low, and they fail to consider the combined characteristics of acoustic and vibration signals. Summary of the Invention

[0007] The object of the present invention is to provide a method and system for real-time prediction of lithology based on acoustic vibration signal characteristics during drilling and rock crushing.

[0008] The present invention provides a method for constructing a lithology model for real-time prediction based on acoustic and vibration signal characteristics during drilling and rock fragmentation, which comprises the following steps:

[0009] 1) obtaining a sound signal and a vibration signal of rock breaking during drilling, wherein the vibration signal includes an x-axis vibration signal, a y-axis vibration signal, and a z-axis vibration signal;

[0010] 2) preprocessing the aforementioned sound signal and vibration signal, including pre-emphasis, wavelet threshold noise reduction, bandpass filtering and time-frequency domain analysis, to obtain time-frequency images of the sound signal, x-axis vibration signal, y-axis vibration signal and z-axis vibration signal;

[0011] 3) Based on AlexNet, the CBAM attention mechanism and residual structure were added to build the AlexNet-RCBAM convolutional neural network. The time-frequency image obtained in step 2) was used as the dataset for data augmentation. The dataset was split and input into the AlexNet-RCBAM convolutional neural network for training and optimization to build a basic model for four signals: sound, x-axis vibration, y-axis vibration, and z-axis vibration.

[0012] 4) Integrate the four basic models in step 3) to obtain an integrated model.

[0013] In step 2), the pre-emphasis is to filter the acoustic vibration signal with a first-order FIR high-pass filter, and the transfer function of the filter in the z domain is as follows:

[0014] H(z)=1-az -1 (Formula 1)

[0015] Where a is the pre-emphasis coefficient, which is usually in the range of 0.9<a<1, and preferably a=0.95;

[0016] And / or, in step 2), the wavelet threshold denoising selects the Sym8 wavelet basis function for filtering and denoising, the decomposition layer number is 9, and the threshold quantization method processes the wavelet coefficients with a soft-hard compromise threshold function, and the formula is as follows:

[0017]

[0018] Where w j is the wavelet coefficient of the jth layer; w jn is the wavelet coefficient after the j-th layer threshold function processing; α∈[0,1], take α=0.5; T is the threshold, and the calculation formula is as follows:

[0019]

[0020] Where w1 is the wavelet coefficient of the first layer; N is the length of the signal;

[0021] And / or, in step 2), the bandpass filtering process is to filter the acoustic vibration signal using a sixth-order Butterworth bandpass filter, and its z-domain transfer function is as follows:

[0022]

[0023] Where a(m) and b(n) are coefficient vectors; N is the order of the bandpass filter;

[0024] Among them, the lower limit frequency of the band-pass filter for processing sound signals is 100Hz and the upper limit frequency is 20000Hz; the lower limit frequency of the band-pass filter for processing vibration signals is 3000Hz and the upper limit frequency is 10000Hz;

[0025] And / or, in step 2), the time-frequency domain analysis uses STFT to perform time-frequency analysis: during the analysis, a Hamming window w(n) with a window length of NFFT is selected to intercept the original signal x(n), and the windowed signal is Fourier transformed. After that, the window is moved by the length of the Hop Size each time, and the two adjacent windows are overlapped by the length of the Overlap. After windowing, the Fourier transform is continued, and finally a two-dimensional matrix S(ω, τ) is obtained. The calculation formula is as follows:

[0026]

[0027] Since the acoustic vibration signal is a power signal, its energy should be measured by the power spectrum density. Therefore, the amplitude S(ω, τ) at the corresponding time and frequency should be converted into the power spectrum density P(ω, τ). The calculation formula is as follows:

[0028]

[0029] Where, f r is the frequency resolution; f s is the sampling rate; N FFT is the window length of the Hamming window;

[0030] Preferably, the analysis window length N FFT 512, 1024 or 2048 respectively.

[0031] In step 3), the data enhancement is performed in sequence using three methods: color dithering, noise perturbation, and blurring.

[0032] In step 3), the training and tuning are based on the Tensorflow architecture; the tuning is to tune the hyperparameters, and the adjustment objects include weight initialization method, learning rate decay strategy, learning rate decay factor, learning rate and number of iterations.

[0033] In step 4), the integration is performed using the weighted average method, the weighting method is to weight the output class probabilities of the base classifiers, and the weights are assigned using the inverse mean square method;

[0034] Preferably, the vector output by the AlexNet-RCBAM network is processed by the Softmax function and converted into the probability corresponding to each category, as shown in the following formula:

[0035]

[0036] Where n is the number of recognition task objects; a i is the i-th element of the vector input to Softmax; p i The probability that the model identifies the data as class i;

[0037] The weighted harmonic mean F R The calculation formula for the precision and recall rate of the comprehensive base classifier is as follows:

[0038]

[0039] Where P is the precision rate; R is the recall rate; β is the weight between P and R.

[0040] F R The value represents the ability of the base classifier to accurately identify a certain type of lithology, so 1-F R The value of can represent the prediction error relative to perfect recognition. Based on the principle that a large error should be assigned a small weight and a small error should be assigned a large weight, the weight is assigned using the inverse mean square method. The calculation formula is as follows:

[0041]

[0042] Where, F βij W is the weighted harmonic mean of the j-th base classifier identifying the i-th class object; ij The weight assigned to the j-th base classifier when identifying the i-th class object, and satisfying

[0043] After integration, the probability that the strong classifier identifies the data as the i-th category is P wi =Wi1 p i1 +W i2 p i2 +…+W im p im ;

[0044] More preferably, β is selected from 1.0 or 1.5.

[0045] The present invention also provides a real-time lithology prediction model based on acoustic vibration signal characteristics during drilling and rock crushing, which is constructed by the method.

[0046] The present invention also provides a system for real-time prediction of lithology based on acoustic vibration signal characteristics during drilling and rock fragmentation, which comprises:

[0047] An input module is used to input a sound signal and a vibration signal, wherein the vibration signal includes an x-axis vibration signal, a y-axis vibration signal, and a z-axis vibration signal;

[0048] A calculation module, which calculates the final prediction result of the rock mass characteristics according to the model module constructed by implementing any one of the methods of claims 1 to 5 and the information of the input module;

[0049] Output module, used to output the final rock mass characteristic prediction results.

[0050] The present invention also provides a method for real-time prediction of lithology based on acoustic vibration signal characteristics during drilling and rock crushing, the steps of which are as follows:

[0051] (1) obtaining a sound signal and a vibration signal of the drilled rock to be inspected, wherein the vibration signal includes an x-axis vibration signal, a y-axis vibration signal, and a z-axis vibration signal;

[0052] (2) preprocessing the aforementioned sound signal and vibration signal respectively, wherein the preprocessing includes pre-emphasis, wavelet threshold noise reduction, bandpass filtering and time-frequency domain analysis to obtain time-frequency images of the sound signal, x-axis vibration signal, y-axis vibration signal and z-axis vibration signal;

[0053] (3) Input the time-frequency image obtained in step 2) into the model constructed by any one of the methods of claims 1 to 5, and draw a conclusion by comparison.

[0054] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program is used to implement the method for real-time prediction of lithology based on acoustic vibration signal characteristics during drilling and rock crushing as described in claim 9.

[0055] The experimental results show that the present invention collects the acquisition method of acoustic vibration signals during drilling and crushing rock in real time, as well as the basic information of seven types of rock properties. Then, the influence of rotation speed, drilling speed, drill bit diameter and drilling depth on the test equipment and drilling process is analyzed. To improve the generalization of subsequent models, seven types of rocks are drilled and acoustic vibration signals are collected under a comprehensive test with three factors and three levels.

[0056] The present invention pre-emphasizes the measured signal to compensate for energy loss in the high-frequency band, performs wavelet threshold noise reduction and bandpass filtering to filter out equipment internal noise and environmental noise, reduces the serious energy loss in the high-frequency band of the acoustic vibration signal when it propagates in the air and rock, and during the instrument acquisition and transmission process, and ultimately reduces the presence of environmental noise, circuit defects, and noise generated by signal transmission in the measured signal, thereby improving the reliability and accuracy of signal processing and analysis.

[0057] This method systematically analyzes the lithologic and drilling characteristics of acoustic vibration signals. Focusing on seven types of lithology, this method designs comprehensive, multi-factor, and multi-level experiments to collect acoustic vibration signals under various conditions. This analysis uses time-domain characteristic parameters to analyze the lithologic and drilling characteristics in the time domain. This analysis is further analyzed in the frequency domain using information such as power spectral density. Finally, time-frequency analysis is used to generate rich, feature-rich time-frequency images to meet the needs of model construction.

[0058] The present invention constructs base classifiers for each signal direction and combines them into an integrated model with high accuracy and strong generalization. This method builds a convolutional neural network for lithologic identification. Using deep learning, the method constructs base classifiers for each signal direction and interprets the decision basis. After determining the base classifier combination strategy, the method uses ensemble learning to integrate the base classifiers' decisions to achieve complementary lithologic features, effectively improving the accuracy and generalization of the integrated model.

[0059] Obviously, based on the above contents of the present invention, according to common technical knowledge and customary means in this field, without departing from the above basic technical ideas of the present invention, other various forms of modifications, replacements or changes can be made.

[0060] The following further describes the above content of the present invention in detail through specific embodiments in the form of examples. However, this should not be construed as limiting the scope of the above subject matter of the present invention to the following examples. All technologies implemented based on the above content of the present invention fall within the scope of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 Acoustic vibration signal acquisition system for drilling rock

[0062] Figure 2 Sound focusing device

[0063] Figure 3For drilling crushed rock samples

[0064] Figure 4 The main role of influencing factors

[0065] Figure 5 Amplitude-frequency and phase-frequency characteristics of first-order FIR high-pass filter

[0066] Figure 6 Time domain acoustic vibration signal before and after pre-emphasis

[0067] Figure 7 Time domain acoustic vibration signal before and after wavelet threshold denoising

[0068] Figure 8 Amplitude-frequency and phase-frequency characteristics of Butterworth bandpass filter

[0069] Figure 9 Time domain acoustic vibration signal before and after bandpass filtering

[0070] Figure 10 Time-frequency analysis process

[0071] Figure 11 Time-frequency diagrams under different window lengths

[0072] Figure 12 3D and 2D time-frequency images

[0073] Figure 13 Time-frequency diagram of acoustic vibration signals for seven types of lithology (Note: Drilling conditions are: drill bit diameter 30mm, rotation speed 400RPM, drilling speed 3mm / min, drilling depth 20mm)

[0074] Figure 14 Artificial neural network structure diagram

[0075] Figure 15 Convolutional Neural Network Architecture

[0076] Figure 16 Convolutional layer and pooling layer

[0077] Figure 17 Fully connected layer

[0078] Figure 18 CBAM attention mechanism

[0079] Figure 19 Residual structure

[0080] Figure 20 AlexNet-RCBAM network structure

[0081] Figure 21 Lithology identification model construction process

[0082] Figure 22Time-frequency image data enhancement: (a) Time-frequency image of granite sound; (b) color dithering; (c) noise perturbation; (d) blurring

[0083] Figure 23 Changes in loss and accuracy for model training under various conditions: (a) learning rate 1×10-4, Batch_size 256; (b) learning rate 1×10-4, Batch_size 128; (c) learning rate 1×10-4, Batch_size 64; (d) learning rate 5×10-5, Batch_size 256; (e) learning rate 5×10-5, Batch_size 128; (f) learning rate 5×10-5, Batch_size 64

[0084] Figure 24 The loss and accuracy changes of the model when training on the vibration dataset: (a) training based on the x-axis dataset; (b) training based on the y-axis dataset; (c) training based on the z-axis dataset

[0085] Figure 25 Confusion matrix of the test set predicted by the corresponding model: (a) Model recognition effect based on sound signal; (b) Model recognition effect based on x-axis vibration signal; (c) Model recognition effect based on y-axis vibration signal; (d) Model recognition effect based on z-axis vibration signal

[0086] Figure 26 EM-AlexNet-RCBAM integrated model integrating probability weighting mechanism

[0087] Figure 27 Performance of the EM-AlexNet-RCBAM ensemble model at different β values

[0088] Figure 28 Borehole lithology prediction DETAILED DESCRIPTION

[0089] The raw materials and equipment used in the present invention are all known products and are obtained by purchasing commercially available products.

[0090] Example 1

[0091] 1. Acquisition and preprocessing of acoustic vibration signals

[0092] (1) Acoustic vibration signal acquisition system

[0093] The present invention uses a micro-drilling test bench for rock drilling, whose model is a ZK5040BS / 2 vertical CNC drilling rig, to analyze the influence of drilling factors on model performance.

[0094] During the test, the micro-drilling test bench's motor drives the drill rod to rotate at a set speed. This rotation drives the diamond-impregnated drill bit at the bottom. Simultaneously, the drill bit, driven by a feed mechanism, cuts into the rock, drilling in a rotary motion to the desired depth. Throughout the drilling process, a cooling system nozzle sprays clean water to cool the drill bit and remove rock cuttings from the hole. As the drill bit breaks rock, vibration and sound signals are collected by pre-installed accelerometers and microphones, respectively. The data acquisition card's A / D converter converts the analog signals into digital signals, which are then stored in a host computer.

[0095] Drilling rock vibration signal acquisition system Figure 1 shown.

[0096] Rock vibration during drilling exhibits spatial randomness. To fully capture the rock-crushing vibration signal, a triaxial accelerometer with a range of 50g, a sensitivity of 95.528mV / g, and a response frequency of 20 to 10,000Hz was used. The drill bit rotated along the x- and y-axes, while the drill rod advanced along the z-axis. To capture the rock-crushing acoustic signal, a microphone with a dynamic range of 16 to 140dB, a sensitivity of 51.28mV / Pa, and a response frequency of 10 to 20,000Hz was selected. The microphone was also equipped with a preamplifier for impedance conversion and preamplification. Because acoustic signals are susceptible to noise interference, a sound focusing device was designed in addition to the microphone and preamplifier. Sound wave reflection follows the law of reflection. When the incident sound wave contacts the concave parabola interface, the sound wave is reflected and forms an acoustic focus. Therefore, when collecting signals, the sound focusing device is attached to the rock surface so that the sound signal is focused inside the device and collected by the microphone. This collection method can reduce the interference of environmental noise to improve the signal-to-noise ratio. The sound focusing device is as follows: Figure 2 shown.

[0097] The data acquisition card connected to the sensors has four channels, each used to receive analog signals from the four sensor cables. Based on the Nyquist sampling theorem, the sampling frequency of the data acquisition card is set to 51200 Hz.

[0098] (2) Basic information of rocks

[0099] The research on the lithology intelligent identification method of the present invention takes seven types of lithology, namely marble, marl, limestone, quartz sandstone, granite, feldspathic sandstone and shale, as the identification objects. The information of each type of lithology is shown in Table 1.

[0100] Table 1 Basic information of various lithologies

[0101]

[0102] The seven lithology categories, based on their origins, include sedimentary rocks such as marl, limestone, quartz sandstone, feldspathic sandstone, and shale, as well as igneous and metamorphic rocks such as granite and marble. Some of these lithologies share common mineralogy types, while others exhibit significant differences in mineral composition. Besides the primary minerals, rock mechanical parameters also vary among lithologies. Marl, limestone, and shale possess the highest compressive and tensile strengths, with values relatively close to each other. Marble and granite have intermediate compressive and tensile strengths, with values very close to each other. Quartz sandstone and feldspathic sandstone, however, exhibit significantly lower compressive and tensile strengths. Therefore, these seven lithology categories share both similarities and differences, providing valuable insights into the development of intelligent lithology identification methods.

[0103] In order to make the micro-drill test bench firmly clamp the sample and reduce relative vibration, the size of each type of rock sample is processed into 20cm×20cm×10cm, such as Figure 3 shown.

[0104] (3) Drilling test design

[0105] The interaction between the drill bit and rock is a complex mechanical process, and sound and vibration are the physical information generated by this process. Therefore, the acoustic and vibroacoustic signal is a carrier of comprehensive information. When collecting acoustic and vibroacoustic signals, it is necessary to consider the differences caused by different lithologies and the influence of factors such as drilling parameters to ensure the generalization ability of the lithology identification model proposed in this paper.

[0106] In order to improve the generalization of the model, based on the micro-drilling conditions, the present invention chooses to add three factors: rotation speed, drilling speed and drill bit diameter for exploration.

[0107] In addition to the three drilling parameters of rotation speed, drilling speed and drill bit diameter, the drilling depth can also be taken into consideration because the homogeneity of each type of rock is different, the distribution of minerals is also different, and the acquisition conditions will be affected as the distance between the drill bit and the sensor changes. The influence of various factors such as Figure 4 .

[0108] Three levels were set for each factor, including drill bit diameter, rotational speed, and drilling rate. In a comprehensive 3×3×3 test, each rock type was drilled and corresponding acoustic vibration signals were collected. Drilling depth increases over time as the drill bit breaks the rock, so there's no need to specifically set different levels for this factor. However, to capture representative acoustic vibration signals during bit penetration and stable drilling, a final drilling depth of 20 mm was set based on the 10 mm drill bit height used in the test. Subsequent feature analysis compared acoustic vibration signals at different hole depths. The drilling test design is shown in Table 2.

[0109] Table 2 Level design of each factor for each lithology

[0110]

[0111] 27 test groups were designed for each type of lithology. Collecting acoustic vibration signals in multiple groups can not only effectively increase the samples required for model construction, but the diversity of the data can also improve the generalization of the model.

[0112] (4) Signal filtering and noise reduction (preprocessing)

[0113] According to the above steps, the acoustic vibration signals, specifically the sound signal and the vibration signal, are obtained. The vibration signals include the x-axis vibration signal, the y-axis vibration signal and the z-axis vibration signal, and are preprocessed respectively. The preprocessing is specifically carried out by pre-emphasis, wavelet threshold noise reduction, bandpass filtering and time-frequency domain analysis in sequence.

[0114] 1. Pre-emphasis

[0115] Pre-emphasis has no effect on noise, so compensating for high-frequency components can effectively improve the signal-to-noise ratio. Pre-emphasis can be used to filter the acoustic vibration signal using a first-order FIR high-pass filter. The transfer function of the filter in the z domain is as follows:

[0116] H(z)=1-az -1 (Formula 1)

[0117] Where a is the pre-emphasis coefficient, which is usually in the range of 0.9<a<1. In order to facilitate the subsequent analysis of the characteristic information of the acoustic vibration signal, a=0.95 is taken. The amplitude-frequency characteristic and phase-frequency characteristic of this filter are as follows: Figure 5 As shown in the figure, the effect of a time domain acoustic vibration signal before and after the filter processing is as follows: Figure 6 shown.

[0118] 2. Wavelet threshold denoising

[0119] The wavelet threshold denoising method of the present invention selects Sym8 wavelet basis function for filtering and denoising of acoustic vibration signals, and the number of decomposition layers is 9;

[0120] In the threshold quantization method, the wavelet coefficients are processed with a soft-hard compromise threshold function, and the formula is as follows:

[0121]

[0122] Where w j is the wavelet coefficient of the jth layer; w jn is the wavelet coefficient after the j-th layer threshold function processing; α∈[0,1], take α=0.5; T is the threshold, and the calculation formula is as follows:

[0123]

[0124] Where w1 is the wavelet coefficient of the first layer; N is the length of the signal.

[0125] The wavelet coefficients of each layer are processed by threshold to reconstruct the signal. The time domain acoustic vibration signal before and after wavelet threshold denoising is as follows: Figure 7 shown.

[0126] 3. Bandpass filtering

[0127] The present invention designs a band-pass filter with a lower limit frequency of 100 Hz and an upper limit frequency of 20,000 Hz and a band-pass filter with a lower limit frequency of 3,000 Hz and an upper limit frequency of 10,000 Hz to process sound signals and vibration signals respectively.

[0128] The acoustic vibration signal is filtered by a sixth-order Butterworth bandpass filter, and its z-domain transfer function is as follows:

[0129]

[0130] Where a(m) and b(n) are coefficient vectors; N is the order of the bandpass filter.

[0131] The amplitude-frequency characteristics and phase-frequency characteristics of the two bandpass filters are as follows: Figure 8 As shown in the figure, the time domain acoustic vibration signal before and after bandpass filtering is as follows: Figure 9 shown.

[0132] 4. Time-frequency domain analysis

[0133] The present invention describes the signal in a manner of combining three physical quantities of time, frequency and power, namely time-frequency analysis.

[0134] 1) Time-frequency analysis method

[0135] The present invention adopts short-time Fourier transform (STFT) to perform time-frequency analysis and preferably selects a time-frequency diagram with better resolution.

[0136] The time-frequency analysis process of the present invention is as follows Figure 10 , the present invention selects the window length as N FFT The Hamming window w(n) intercepts the original signal x(n) and performs Fourier transform on the windowed signal. Then, the window is moved by the length of the Hop Size each time, and the two adjacent windows are overlapped by the length of the Overlap. The Fourier transform is continued after the windowing. Finally, a two-dimensional matrix s(ω, τ) is obtained to reduce spectrum leakage and Gibbs phenomenon. Each column vector represents the spectrum under the corresponding window. The calculation formula is as follows:

[0137]

[0138] Since the acoustic vibration signal is a power signal, its energy should be measured by the power spectrum density. Therefore, the amplitude S(ω, τ) at the corresponding time and frequency should be converted into the power spectrum density p(ω, τ). The calculation formula is as follows:

[0139]

[0140] Where, f r is the frequency resolution; f s is the sampling rate; N FFT The length of the Hamming window.

[0141] Since STFT has the problem of difficulty in taking both time resolution and frequency resolution into account at the same time, in order to meet the needs of lithology identification, the comparative analysis window length N FFT The display effects of the time-frequency diagrams obtained when the values are 512, 1024, and 2048 are as follows: Figure 11 As shown in the figure, when N FFT When N is 512, the frequency resolution is low due to the small window size, and stripes along the frequency direction appear on the time-frequency graph, and the overall image display effect is torn; when N is FFT When N is 2048, the time resolution is low due to the large window, the time-frequency graph shows pixel grids along the time direction, and the overall image display effect is blurred. FFT The power spectrum density at the corresponding frequency and time on the time-frequency diagram of 1024 is clearly displayed, and its time resolution and frequency resolution have good comprehensive performance.

[0142] 2) Time-frequency image of acoustic vibration signal

[0143] The time-frequency diagram can be visually interpreted as the result of projecting the three-dimensional space representing time, frequency and power spectrum density onto the horizontal plane along the vertical direction, and the color depth represents the value of the power spectrum density, such as Figure 12 The time-frequency diagram shows the relationship between frequency and time from the perspective of the signal's time-frequency domain, and also displays the power spectrum density at the corresponding frequency and time, which can carry information in both the time domain and the frequency domain.

[0144] As can be seen from the figure, the color change of the time-frequency diagram along the frequency direction is completely consistent with the curve change of the power spectrum density diagram obtained by frequency domain analysis, and the color change along the time direction is also very consistent with the trend of the time domain signal waveform. Therefore, the feature extraction of the time-frequency diagram in the present invention can more completely reflect the drilling status and lithology.

[0145] Figure 13 The following is a time-frequency diagram of the acoustic vibration signals for seven types of lithology under specific drilling conditions. The time-frequency diagram shows that the dominant frequency distribution and power levels of the seven types of lithology acoustic vibration signals vary, making them suitable for lithology identification. However, the time-frequency diagram contains complex feature information, making manual extraction difficult to fully utilize the data's characteristics and inherently subjective. Using deep learning to process unstructured image data can better accomplish lithology identification.

[0146] 2. Research on lithology identification model based on deep learning

[0147] (1) Building a neural network model

[0148] 1. Basic principles of deep learning

[0149] Deep learning is a special form of neural network, and its basic model framework is based on artificial neural network (ANN). ANN is inspired by the structure of the human brain, simulates the biological neural system's processing mechanism of complex information, and has powerful learning capabilities. The structure of ANN is as follows: Figure 14 As shown in the figure, it is a highly complex nonlinear model consisting of a large number of interconnected neurons, which is structurally divided into input layer, hidden layer and output layer.

[0150] ANN transmits information from the previous layer to the next layer through forward propagation. Each node in the figure is a neuron, and the connection between two neurons represents the weight. After the information is input into the ANN, each neuron receives stimulation from the upper layer neurons, and the weight ω n After determining the stimulus intensity and bias b to determine the activation ability, the activation function f unction Output information Z and continue to pass it to the next layer. The calculation formula is as follows:

[0151] Z=function(ω1a1+ω2a2+…+ω n a n +b) (Formula 7)

[0152] After information is transmitted through multiple hidden layer neurons, the information features are fully extracted, and the output layer provides the final prediction. Because the neural network is often initialized with random weights during network training, there is a certain difference between the final prediction given by the ANN and the actual situation. The size of the difference is measured by the loss function. The present invention propagates the error backward through the backpropagation method, and calculates the partial derivatives of the neurons layer by layer to adjust the weights so that the loss function converges to a local minimum. Therefore, the ANN learns features through forward propagation to make predictions, and adjusts the network weights through backpropagation to improve the model accuracy, thereby minimizing the average loss function over the entire training set.

[0153] ANN decomposes features with hidden layer neurons and represents higher-level complex features by combining low-level features. It has excellent learning ability and intelligence.

[0154] 2. Convolutional Neural Network Structure

[0155] Convolutional Neural Networks (CNN) is one of the representative algorithms of deep learning that has excellent performance in image processing tasks. The structure of CNN also consists of input layer, hidden layer and output layer, where the hidden layer includes convolution layer, pooling layer and fully connected layer, such as Figure 15 shown.

[0156] Classic neural networks transmit information based on a fully connected model, meaning that any neuron in the next layer is connected to every neuron in the previous layer, with the corresponding weights determining the strength of the connection. When the network structure is complex, the number of connection parameters in a classic neural network can become very large, leading to problems such as slow training, loss of spatial information, and network overfitting. Based on the fact that human visual cognition proceeds from the local to the global, CNNs extract local information through the local receptive fields of convolutional and pooling layers, and use shared weights and biases during convolution operations, effectively reducing the number of network parameters.

[0157] The information extraction process of CNN is as follows Figure 16 As shown in the figure, the local receptive field of the input data is multiplied by the convolution kernel and activated by the activation function to obtain a value. Then, its position is moved by a specified number of steps and the convolution and activation are continued to obtain an output feature map B composed of multiple values. The entire convolution process realizes the feature mapping from the input data layer A to the next layer B. The mapping weight ω and bias b are shared weights and shared biases. The convolution process is calculated as follows:

[0158]

[0159] The size of the output feature map after feature mapping is calculated as follows:

[0160]

[0161] Where As is 0 or 1, representing the length and width directions respectively; S is the number of receptive field movement steps; D is the input data size; K is the convolution kernel size; Pad is the input data padding size; and D′ is the output feature map size.

[0162] To reduce overfitting and improve model robustness, the convolutional feature map B needs to be pooled to reduce data dimensionality and the number of parameters. Pooling methods are mainly divided into average pooling and maximum pooling. During the processing, the local receptive field is used as the object, and the data within the range is moved a specified number of steps in sequence. The average value or the maximum value is calculated, and the output feature map C after feature dimensionality reduction is finally obtained.

[0163] The local information extracted by the convolutional layer and the pooling layer will be integrated into the overall information at the fully connected layer, such as Figure 17As shown in the figure, the fully connected layer uses a fully connected approach for feature mapping. Each neuron in the fully connected layer is connected to all neurons in the previous layer, and each connection has a corresponding weight. Therefore, the fully connected layer has the most parameters in the entire network. The feature information output by neurons in the fully connected layer is highly repetitive, making the model prone to overfitting during training. The random inactivation function dropout can disable the activation function of some neurons (as shown in the dotted line in the figure), thereby reducing repetitiveness between neurons and improving the generalization of the model.

[0164] CNN learns features through forward propagation, calculates the loss function at the network output layer, and backpropagates the error to calculate the partial derivatives to update the weights, thereby enabling the network to train the model to fit in the process of feature learning and weight adjustment.

[0165] 3. Network construction

[0166] When training a CNN model for lithologic identification, the hidden layer neurons extract the characteristic information contained in the time-frequency images. Therefore, the neural network's expressive power directly impacts the model's accuracy and generalization. To improve accuracy and generalization, accelerate training, and prevent overfitting, we chose to base our CNN model on AlexNet and incorporate the CBAM attention mechanism and residual structure.

[0167] 1) CBAM Attention Mechanism

[0168] In the human visual attention mechanism, vision locks onto local targets within the global image, devoting more attention resources to capturing detailed information. This focus on the target while ignoring other information effectively utilizes limited attention resources and improves the human brain's ability to analyze visual information. In the field of deep learning, the expressive power of a neural network depends on the number of neurons in the hidden layer. The more hidden neurons, the stronger the neural network's expressive power. Consequently, the model needs to store more information, which can easily lead to information overload. However, the importance of network input data varies. By incorporating an attention mechanism, the neural network can focus on key information in the data, thereby addressing information overload, improving network computational efficiency, and enhancing the model's expressive power.

[0169] The attention mechanism CBAM can be added to any convolutional layer and has extremely high portability. Its structure consists of a channel attention module (CAM) and a spatial attention module (SAM). The specific structure of the CBAM added to the lithology recognition model training network is as follows: Figure 18 .

[0170] The feature map M output by the convolutional layer of the neural network is first analyzed by CAM to determine the distribution of important features along the channel dimension. During this process, the feature map M, which is sized H×W×C, is pooled into two 1×1×C feature maps along the height H and width W directions using Maxpool and Avgpool, respectively. The feature maps are then fed into a fully connected neural network (MLP) constructed with a depth of two layers and C / 4 and C neurons. The channel weights of the feature maps are analyzed, and the two output feature maps are added together and activated by function F to obtain Wc. Wc is then multiplied by the original feature map M to obtain the CAM output feature map M'. The calculation process is as follows:

[0171] Wc=F(MLP(Maxpool(M))+MLP(Avgpool(M))) (Formula 10)

[0172]

[0173] The CAM output feature map M' is then analyzed by SAM for the spatial distribution of important features. The feature map M' is pooled along channel C using Maxpool and Avapool to form two H×W×1 feature maps. This is then concatenated into an H×W×2 feature map using the Concat method. This map is then reduced in dimension using a 7×7 convolution kernel and activated to produce Ws. Ws is then multiplied with the feature map M' to produce the CBAM output feature map M", calculated as follows:

[0174] Ws=F(f 7×7 ([Maxpool(M'); Avgpool(M')])) (Formula 12)

[0175]

[0176] 2) Residual structure

[0177] In theory, deeper neural networks have better expressive power. However, in practice, deeper networks become more difficult to train and more prone to network degradation. Furthermore, excessively deep networks can lead to problems such as vanishing and exploding gradients. This occurs when the error backpropagates, making it difficult to update the weights of layers near the neural network input or causing the new weights to become too large. To address these issues, layers that create an identity mapping (where input and output are equal) can be added to the neural network. However, in practice, this is not achieved directly by training the layers to the identity mapping, but rather indirectly through a residual structure.

[0178] The residual structure set for training the lithology recognition model is as follows: Figure 19The input feature map M undergoes two layers of 3×3 convolution processing to obtain the feature map M'. The original feature map M is directly added to the feature map M' through a shortcut connection and activated, finally obtaining the output feature map M". The analysis process of the residual structure can be reflected as M" = M + M'. During this process, the residual function M' = M" - M continuously learns and approaches 0, ultimately achieving the identity mapping between the input feature map M and the output feature map M". Overall, the residual structure not only deepens the depth of the neural network to improve the model's expressiveness, but also trains the network layer to the identity mapping, preventing problems such as network degradation, gradient vanishing, and gradient explosion.

[0179] 3) AlexNet-RCBAM network structure

[0180] AlexNet is a deep convolutional neural network proposed in 2012. Its emergence is an important milestone in the development of computer vision. AlexNet uses the Relu activation function to speed up network training and prevent gradient disappearance, uses LRN normalization to improve the generalization ability of the model, and finally uses Dropt in the fully connected layer to avoid model overfitting, with excellent performance. Taking into account the characteristics of the complex time-frequency image information of the acoustic vibration signal and the large demand for model generalization ability, when building the network, it is chosen to add the CBAM attention mechanism and residual structure on the basis of AlexNet to solve problems such as information overload, gradient disappearance and gradient explosion, while improving the network depth. In addition, since the graphics card GeForce RTX 2060 used in the present invention meets the resource requirements of network training, a network structure based on a single GPU is designed. The constructed AlexNet-RCBAM network structure is as follows Figure 20 .

[0181] As shown in the figure, before the original time-frequency image is input to the network, the B, G, and R channels are extracted and resized to form a 227×227×3 input image. AlexNet-RCBAM has seven convolutional layers, each with a different kernel size, and the last three layers have a residual architecture. There are three pooling layers, all using the Maxpool method with a 3×3 kernel size. There is one CBAM module, with the CAM module containing two MLP layers and the SAM module containing one 7×7 convolution operation. There are three fully connected layers, with the first two layers having a size of 4096 and the last layer having a size of 7. The output vector after Softmax processing represents the probability that the time-frequency image belongs to each rock type. The detailed configuration of each AlexNet-RCBAM layer is shown in Table 3.

[0182] Table 3 AlexNet-RCBAM network detailed configuration

[0183]

[0184] As shown in the table, AlexNet-RCBAM has a total of 15 layers, 12 of which involve parameter calculation. The network has a total of 59,556,457 parameters related to weights and biases, of which 91.61% come from the fully connected layers. Compared to before the addition of the residual structure and CBAM attention mechanism, the network's parameter count increased by only 2.14%. Training the AlexNet-RCBAM model requires 454MB of video memory due to simultaneous forward and backward propagation. The amount of video memory used for output depends on the batch size (Batch_size). Upon completion, the model occupies approximately 227MB of memory. Therefore, AlexNet-RCBAM offers strong expressive power and requires relatively few computer resources.

[0185] In summary, the CBAM attention mechanism and residual structure are added to AlexNet to build the AlexNet-RCBAM convolutional neural network.

[0186] (2) Model construction

[0187] 1. Model building process

[0188] The time-frequency images of the sound signal, x-axis vibration signal, y-axis vibration signal and z-axis vibration signal obtained in step 1 are used to construct a model.

[0189] The lithology identification model is constructed based on the deep learning framework Tensorflow-gpu 2.5.0. The operating system used is Windows, the CPU model is i5-9300H, and the GPU model is GeForce RTX 2060. The model construction includes four processes: data generation, dataset construction, model training, and model evaluation. The detailed process is as follows Figure 21 shown.

[0190] 2. Dataset Construction

[0191] 1) Data augmentation:

[0192] After filtering and noise reduction, and time-frequency analysis of the acoustic and vibroacoustic signals, a total of 3,780 time-frequency images representing seven lithologic types were obtained for each of the four data types: sound, x-axis vibration, y-axis vibration, and z-axis vibration. Unlike photographic images, time-frequency images are generated through digital signal processing and are immune to image content flipping or offset. However, image color distribution and clarity may be affected by factors such as detection methods, instrument performance, and analysis techniques. Therefore, pixel transformation was used for data enhancement.

[0193] The time-frequency image is enhanced by three methods: color jitter, noise disturbance and blur processing. The effects before and after data enhancement are as follows: Figure 22When performing color dithering, the image's saturation, brightness, or contrast are randomly adjusted, affecting the image's brightness, vividness, and contrast, respectively. When performing noise perturbation, Gaussian noise with a mean of 0 and a variance of 0.03 is added to achieve noise addition without changing the original image's brightness. Blur processing uses a median filter with a filter kernel size of 9×9. During processing, the grayscale values at the intersection of the filter kernel and the pixel are sorted, and the middle value is assigned to the pixel center. The filter kernel is then continuously moved and assigned values to blur the image. After data enhancement, the time-frequency images of the four types of data each had a dataset size of 15,120, four times the size of the original dataset, significantly expanding the dataset.

[0194] 2) Data Splitting:

[0195] Each dataset is divided into a training set, a validation set, and a test set for three purposes: model training, hyperparameter tuning, and model evaluation. Because the base classifiers constructed on each dataset will be subsequently integrated into a strong classifier, the test set is further divided into a base classifier test set and a strong classifier test set, which are used to evaluate the performance of individual models and the performance of the integrated model, respectively. Furthermore, to ensure that the four time-frequency images input to the strong classifier are derived from the same acoustic vibration signal at the same moment, the strong classifier test sets divided from the four datasets must ensure that the images correspond to each other at the same moment. The training set, validation set, base classifier test set, and strong classifier test set are divided in a ratio of 6:2:1:1. The partitioning strategy for each dataset is shown in Table 4.

[0196] Table 4 Division strategy for each type of data set

[0197]

[0198] 3. Model training and tuning

[0199] Model training is a dynamic process of updating parameters and adjusting hyperparameters to optimize model performance. Model parameters are automatically updated through iteration, while hyperparameters must be manually set before model training. Therefore, during model training, tuning hyperparameters is a necessary step to optimize model performance. Hyperparameters that can be adjusted primarily include weight initialization method, batch size (Batch_size), training period (epoch), learning rate, number of iterations, learning rate decay strategy, and learning rate decay factor.

[0200] Based on the structural characteristics of the neural network, the weight initialization method, learning rate decay strategy, and learning rate decay factor can be determined first. The model weights are continuously updated iteratively based on the initial values, so the weight initialization method affects the gradient propagation and convergence speed of the model. Since the AlexNet-RCBAM activation function is a nonlinear function ReLU, the weight initialization method should choose the He Gaussian distribution proposed by Kaiming He. This can evenly distribute the activation values of each layer output to avoid gradient vanishing. The He Gaussian distribution is as follows:

[0201]

[0202] Where l is the neural network layer with initialized weights; n l is the number of neurons in the l-1 layer.

[0203] Learning rate decay can help the model escape the local optimal solution in the early stages of training and achieve stable convergence in the later stages. Exponential decay is the most commonly used and effective learning rate decay strategy. The decay factor range is 0.96 to 1, and this paper sets it to 0.96. Learning rate decay is based on epochs, and the learning rate of each epoch is as follows:

[0204] α=α0×0.96 Num (Formula 15)

[0205] Where α is the learning rate of each epoch; α0 is the initial learning rate; Num is the number of the current epoch.

[0206] The hyperparameters batch_size, epoch, learning rate, and number of iterations must be determined based on the actual training results of the model. Batch_size is the number of training samples used in each iteration. If the number is too large, the training memory resources required will be large and the model generalization will be poor. If it is too small, the model converges slowly. Epoch is the number of times the entire training set is trained in the model. Too large or too small will lead to overfitting or underfitting of the model. The learning rate is the step size for updating the weight matrix during backpropagation. If the value is too large, the model will have difficulty escaping the local optimal solution. If it is too small, the model converges slowly. The number of iterations is the number of times the model is trained using batch samples. The value is determined by batch_size, epoch, and the size of the training set, as shown in the following formula:

[0207]

[0208] Where iter is the number of iterations; D is the size of the training set.

[0209] According to research experience and the size of the dataset of this invention, the epoch is selected as 50 times; the learning rate is initially selected as 1×10 -4 , 5×10 -5The initial batch_size is 64, 128, and 256. Based on the sound dataset, the model is trained under 6 combinations of hyperparameters. The changes in loss and accuracy during the training process are shown in the following figure. Figure 23 shown.

[0210] As shown in the figure, at the same learning rate, as the batch size decreases, the model converges faster and the validation set accuracy significantly improves. At the same batch size, as the learning rate decreases, the model converges slower and the validation set accuracy decreases. By comparing the loss and accuracy trends of different subgraphs, we determined that 1×10⁻4 is the optimal learning rate and 64 is the optimal batch size. Furthermore, to train the model until it reaches a good fit, we determined that 20 epochs is the optimal number of iterations based on subgraph c, which corresponds to 2835.

[0211] Since the vibration dataset and the sound dataset have the same size and similar data format, under the settings of optimal learning rate and optimal batch_size, the remaining models are trained based on the x-axis vibration dataset, y-axis vibration dataset, and z-axis vibration dataset respectively. The changes in loss value and accuracy are as follows: Figure 24 .

[0212] As can be seen from the figure, when training on the vibration dataset in three directions, the model reaches the optimal value at about the 20th epoch. If training continues, the loss value begins to oscillate and there is no downward trend. Therefore, the optimal epoch when training based on the vibration dataset can also be determined to be 20.

[0213] Finally, the model was retrained based on four types of data sets under the optimal hyperparameter settings. The values of the optimal hyperparameters are shown in Table 5.

[0214] Table 5 Hyperparameter tuning table

[0215]

[0216]

[0217] 4. Model Evaluation Metrics

[0218] The performance of the lithology recognition model is evaluated by using two indicators: precision (P) and recall (R). Precision indicates the proportion of samples identified as a certain lithology type by the model that are actually of that type. Recall indicates the proportion of samples of a certain lithology type that are accurately identified by the model. The calculation formulas for the two indicators are as follows:

[0219]

[0220]

[0221] Where TP is a true positive, which means that the sample is identified as a positive example and is actually a positive example; TN is a true negative, which means that the sample is identified as a negative example and is actually a negative example; FP is a false positive, which means that the sample is identified as a positive example and is actually a negative example; FN is a false negative, which means that the sample is identified as a negative example and is actually a positive example.

[0222] In addition, the confusion matrix can also be used to evaluate the model in detail. The confusion matrix is composed of TP, TN, FP, and FN corresponding to various types of lithology. There are 7 types of lithology required for prediction by the model of the present invention, so it is a 7×7 matrix reflecting the difference between the true value and the predicted value.

[0223] For multi-class lithology identification tasks, the model's ability to identify lithology can be further evaluated using Accuracy, Macro_P, and Macro_R. Among these indicators, Accuracy reflects the proportion of successful predictions when the model identifies lithology, while Macro_P and Macro_R are the average values of the Precision_P and Recall_R for each lithology class, respectively. Each indicator is calculated as follows:

[0224]

[0225]

[0226]

[0227] Where n is the number of lithologies that the model needs to identify.

[0228] 5. Model Performance Evaluation

[0229] The evaluation of the model is based on the test set. The four optimal lithology recognition models as the base classifiers are used to predict their respective base classifier test sets to evaluate the accuracy and generalization of the model. The confusion matrix obtained after normalization of the test results is as follows: Figure 25 shown.

[0230] Depend on Figure 25As can be seen, the four models generally have good recognition performance for various lithologies, but the recognition performance of marble, marl, limestone, and shale on the four models is lower than that of other lithologies. Marble is the worst recognized by the acoustic signal-based model, with 16% and 9% of marble time-frequency images being identified as marl and limestone, respectively. Mud and limestone are more likely to be confused with each other when recognized by the acoustic signal-based model, at 10% and 7%, respectively. Shale recognition is relatively average across the four models, with 28%, 8%, 7%, and 7% of shale time-frequency images being identified as marble, marl, and limestone, respectively. The reasons for the poor ability to distinguish between these four types of lithology can be explained from different perspectives: From a mineralogy perspective, all four rock types contain calcite; from a rock formation perspective, marble is often formed from limestone through metamorphism, while marl is a transitional rock, often found in the transition zone between limestone and claystone; and from a rock strength perspective, limestone, marl, and shale all have large rock mechanical parameters, generating extremely high energy levels and impact characteristics when drilled. Therefore, these possible factors lead to a certain degree of similarity in the characteristics of the acoustic vibration signals generated when drilling into these rocks, which can lead to deviations in the model's recognition results.

[0231] Table 6 shows the precision and recall rates for various rock types using the four models. The macro values and accuracy rates in the table indicate that the four models have different overall performance in identifying lithology. The models based on x-axis and z-axis vibration signals have the strongest overall performance, while the model based on acoustic signals has the worst overall performance. Furthermore, there are significant differences in precision and recall rates among different lithology types. Feldspathic sandstone has the highest overall recall rate across the four models, while quartz sandstone has the highest overall precision rate. Although the acoustic signal-based model has the worst overall performance, quartz sandstone has a precision rate of 0.95 using this model. Overall, the four models exhibit varying levels of precision and recall for each lithology type. Therefore, to comprehensively assess model performance, the four models will be further integrated to leverage their respective strengths and weaknesses.

[0232] Table 6 Model test indicators

[0233]

[0234] From the confusion matrix and performance indicators, it can be seen that the four models have an overall good prediction effect on various lithologies and have strong generalization performance.

[0235] To further measure the performance improvement of the AlexNet-RCBAM model compared to other algorithms, we selected two commonly used traditional machine learning algorithms for image recognition, KNN and SVM, and the AlexNet deep learning algorithm for comparison. After parameter adjustment and optimization, each algorithm model was trained on the corresponding dataset. The final accuracy on the test set is shown in Table 7.

[0236] Table 7 Comparison of model accuracy of various algorithms

[0237]

[0238] Table 7 shows that compared to traditional machine learning models, the deep learning model significantly improves the accuracy of lithology identification, with improvements ranging from 6% to 19%. This indicates that deep learning is more suitable for image recognition tasks than traditional machine learning. Furthermore, by adding the residual structure and CBAM attention machine to AlexNet, the model accuracy increased by 1% to 5%, demonstrating that the AlexNet-RCBAM model further improves its performance in lithology identification compared to AlexNet.

[0239] This paper proposes a lithology identification method based on deep learning of time-frequency images of acoustic vibration signals during drilling and rock crushing. First, the CBAM attention mechanism and residual structure are added to AlexNet to build the AlexNet-RCBAM convolutional neural network for constructing a lithology identification model. The respective datasets of the acoustic vibration signals are then enhanced and divided. After model training and optimization, four lithology identification models are obtained (based on the four signals: sound, x-axis vibration, y-axis vibration, and z-axis vibration). Model evaluation shows that the model has good overall performance, and AlexNet-RCBAM outperforms traditional machine learning algorithms and AlexNet in the task of lithology identification.

[0240] 3. Intelligent lithologic identification method based on multi-model ensemble learning

[0241] 1. Model Ensemble Method

[0242] Traditional voting methods use the majority rule to make the final decision on base classifier selection. This approach assumes that each base classifier model has equal status and ignores the performance differences caused by differences in the datasets used, the algorithms used, and the hyperparameters used during base classifier training. Weighting the output class probabilities based on the performance of the base classifiers can improve the overall performance of the ensemble model. However, in addition to the performance differences between the base classifiers, the recognition ability of a single base classifier for each class also varies. Therefore, weighting cannot be based directly on the base classifiers but should instead be applied to the probabilities of each class.

[0243] The integration of lithology recognition models is based on the output class probability, and the vector output by the AlexNet-RCBAM network is converted into the probability corresponding to each category through the Softmax function, as shown in the following formula:

[0244]

[0245] Where n is the number of recognition task objects; a i is the i-th element of the vector input to Softmax; p i The probability that the model identifies the data as class i.

[0246] Measuring the reliability of each base classifier in identifying a certain type of lithology is the key to realizing the integration of lithology recognition models. Before this, it is necessary to first determine the recognition ability of each base classifier for each type of lithology. In the model performance evaluation, the precision rate and recall rate respectively represent the proportion of samples identified as a certain type of lithology that are actually of this type of lithology and the proportion of samples of a certain type of lithology that are accurately recognized by the model. These two types of indicators have different description angles, but both truly reflect the recognition ability of the base classifier for each type of lithology. Therefore, the present invention uses the weighted harmonic mean F R The calculation formula for the precision and recall rate of the comprehensive base classifier is as follows:

[0247]

[0248] Where P is the precision rate; R is the recall rate; β is the weight between P and R.

[0249] F R The value represents the ability of the base classifier to accurately identify a certain type of lithology, so 1-F R The value of can represent the prediction error relative to perfect recognition. Based on the principle that a large error should be assigned a small weight and a small error should be assigned a large weight, the weight is assigned using the inverse mean square method. The calculation formula is as follows:

[0250]

[0251] Where, F βij W is the weighted harmonic mean of the j-th base classifier identifying the i-th class object; ij The weight assigned to the j-th base classifier when identifying the i-th class object, and satisfying

[0252] After integration, the probability that the strong classifier identifies the data as the i-th category is P wi =W i1 p i1 +W i2 p i2 +…+W im p im, the structure of the EM-AlexNet-RCBAM integrated model is as follows Figure 26 shown.

[0253] 2. Ensemble model selection

[0254] F of each base classifier category R The value will affect the weight of the corresponding category probability, and the β value determines the resulting F R The value of β affects the degree to which the model attaches importance to precision and recall. When β < 1, the model attaches more importance to precision; when β = 1, the model attaches equal importance to both; when β > 1, the model attaches more importance to recall. Therefore, it is necessary to analyze the impact of β on the EM-AlexNet-RCBAM ensemble model. The test results of the ensemble model obtained under different β values on the strong classifier test set are shown below. Figure 27 shown.

[0255] PE represents the performance of the integrated model, which is represented by Figure 27 Analysis shows that the integrated model satisfies PE when different β values are taken β=0.5 <PE β=2.0 <PE β=1.5 =PE β=1.0 , meaning that the performance of the ensemble model first increases and then decreases as the value of β increases. When β is set to 1.0 and 1.5, the performance of the ensemble model differs little. However, to achieve a higher recall rate for the ensemble model and ensure that samples are identified as completely as possible when drilling into a certain lithology, β is ultimately set to 1.5. The class probability weights of the base classifiers and the performance indicators of the ensemble model are shown in Tables 9 and 10, respectively.

[0256] Table 9 Probability weight of each category of base classifier (%)

[0257]

[0258] Table 10 Performance indicators of the integrated model (%)

[0259]

[0260] The accuracy of various models is shown in Table 11. Regardless of the ensemble method used, the accuracy of the model ensemble improved, demonstrating that ensemble learning can effectively improve the overall performance of the model in identifying lithology. Furthermore, among the three ensemble methods, the EM-AlexNet-RCBAM ensemble model achieved the highest accuracy, improving by 1% to 3% compared to the other two methods. Therefore, weighting the output class probabilities of the base classifiers has a greater advantage in ensembles. The ensemble model ultimately improved the accuracy by 5% to 19% compared to a single base classifier, effectively combining the performance of each base classifier and improving the model's generalization.

[0261] Table 11 Accuracy of various models (%)

[0262]

[0263] 3. Generalization Analysis of Integrated Model

[0264] The EM-AlexNet-RCBAM integrated model demonstrated excellent performance on the test set, meeting the basic requirements for lithology identification. However, the construction of the base classifier and the integrated model throughout the process was based on a library of time-frequency images of acoustic vibration signals under existing drilling conditions. Intelligent lithology identification aims to identify lithology to guide the optimization of drilling parameters, so the model needs to be able to accurately identify lithology even after drilling conditions change.

[0265] To further analyze the ability of the integrated model to identify lithology under unknown drilling conditions, a new diamond-impregnated drill bit was replaced, and the seven types of rock samples were re-drilled at an unset rotational speed and drilling rate. In addition, a larger drilling depth was selected for the second test. Detailed information is shown in Table 12.

[0266] Table 12 Drilling conditions of the secondary test

[0267]

[0268] When drilling into the rock, the acoustic vibration signal of the whole process is collected at a sampling frequency of 51200Hz. The acoustic vibration signal is processed and converted into a time-frequency image input into the four channels of the integrated model. The recognition results are arranged in order according to the corresponding hole depth, and the resulting drilling lithology prediction is as follows: Figure 28 shown.

[0269] Figure 28 The white bands in each borehole represent model prediction errors at that location, and the figure also shows the lithology prediction accuracy for the corresponding borehole. As shown in the previous analysis of the acoustic vibration signal characteristics, changing drilling conditions will affect the drilling state characteristics in the acoustic vibration signal, and therefore inevitably affect the accuracy of the model's lithology prediction.

[0270] Depend on Figure 28 The model's prediction accuracy is extremely high for boreholes drilled in marble, granite, and feldspar sandstone, while the prediction performance for limestone is relatively poor. This indicates that changes in drilling conditions have affected the lithology identification task, with varying degrees of impact on different lithologies. However, the model's overall prediction performance remains strong under the new drilling conditions, with an overall accuracy of 92.06%. This demonstrates strong generalization and is capable of completing lithology identification tasks under unknown drilling conditions to a certain extent.

[0271] The experimental results show that the four lithology recognition models have different abilities to identify different lithologies. However, since each model extracts time-frequency image feature information from different angles, the present invention comprehensively analyzes the decision-making of each model. The present invention uses the weighted output class probability of each base classifier as the integration strategy, based on F R The weights of the corresponding categories were calculated using the inverse mean square method, and the impact of the β value on the performance of the ensemble model was analyzed. Finally, a quadratic experiment was conducted to test the performance of the ensemble model under unknown drilling conditions. The results showed that the model performance was minimally affected by the changes in drilling conditions. Overall, the ensemble model has strong generalization capabilities.

[0272] In summary, this paper analyzes the lithologic characteristics and drilling status characteristics contained in the acoustic vibration signal and combines it with a deep learning convolutional neural network to construct a lithologic identification model based on the acoustic vibration signal. The model's decision-making is explained. Furthermore, considering the different lithologic identification capabilities of each model, combined with the idea of joint model decision-making, a multi-model integrated learning intelligent lithologic identification method and system are provided:

[0273] (1) The characteristic information of the acoustic vibration signal was obtained. The stationarity of the acoustic vibration signal was analyzed by using the graph test and the run test, and the characteristic analysis of the acoustic vibration signal was performed at an observation scale of 2s. In the time domain analysis, two types of indicators, dimensional characteristic parameters and dimensionless characteristic parameters, were used to analyze the influence of drilling conditions on energy level and impact characteristics, as well as the performance of each type of rock type in energy level and impact characteristics; in the frequency domain analysis, the influence of drilling conditions on power spectrum density and center of gravity frequency, as well as the relationship between the curve shape and main frequency distribution of rock type and power spectrum density were analyzed. The characteristic analysis in the time domain and frequency domain shows that the characteristic information of the acoustic vibration signal can reflect the rock type and drilling status. Finally, STFT was used to perform time-frequency analysis on the acoustic vibration signal. The time-frequency diagram obtained by the time-frequency analysis contains rich features and is complex unstructured data. It can be used as the data set required for the construction of the rock type identification model.

[0274] (2) A lithology recognition model was constructed and the model was explained. AlexNet-RCBAM was constructed by adding the CBAM attention mechanism and residual structure to the convolutional neural network. After data set partitioning and hyperparameter tuning, four optimal lithology recognition models based on sound signals and triaxial vibration signals were trained respectively. The models were evaluated by confusion matrix and performance index, which showed that their accuracy and generalization were good. The model accuracy rates were 0.78, 0.92, 0.87 and 0.92 respectively, which were 6% to 19% higher than those of KNN and SVM algorithms and 1% to 5% higher than those of AlexNet algorithm. Therefore, the performance of AlexNet-RCBAM was better than that of the other comparison algorithms, effectively improving the model's ability to identify lithology. In the model explanation, Grad-CAM was used to analyze the class activation heat maps of seven types of lithology time-frequency images. The distribution of thermal areas on the heat maps showed a correlation with the rock properties, and the greater the significance of the distribution difference, the better the model performance. In addition, under the same sound dataset size, as the number of rotational speed samples increases, the accuracy of the trained model increases and the distribution range of its thermal area narrows, indicating that increasing drilling conditions has an effect on improving the generalization of the model.

[0275] (3) A multi-model ensemble learning method for intelligent lithology identification was proposed. Based on the idea of joint decision-making by multiple models, the weighted average of the output class probabilities of each category of each model was used as the combination strategy. The prediction error when the β value was 1.5 was used to assign weights using the inverse mean square method. After integration, the accuracy of the EM-AlexNet-RCBAM ensemble model on the test set was 97.43%, which was 5% to 19% higher than that of a single base classifier and 1% to 3% higher than that of other ensemble methods. The ensemble model was used to predict the lithology of the borehole in the secondary test under new drilling conditions, and the accuracy was still 92.06%. The ensemble model has strong generalization ability.

Claims

1. A method for constructing a real-time prediction lithology model based on acoustic and vibration signal characteristics during drilling and rock fragmentation, characterized by: It includes the following steps: 1) obtaining a sound signal and a vibration signal of rock breaking during drilling, wherein the vibration signal includes an x-axis vibration signal, a y-axis vibration signal, and a z-axis vibration signal; 2) preprocessing the aforementioned sound signal and vibration signal, including pre-emphasis, wavelet threshold noise reduction, bandpass filtering and time-frequency domain analysis, to obtain time-frequency images of the sound signal, x-axis vibration signal, y-axis vibration signal and z-axis vibration signal; 3) Based on AlexNet, the CBAM attention mechanism and residual structure were added to build the AlexNet-RCBAM convolutional neural network. The time-frequency image obtained in step 2) was used as the dataset for data augmentation. The dataset was split and input into the AlexNet-RCBAM convolutional neural network for training and optimization to build a basic model for four signals: sound, x-axis vibration, y-axis vibration, and z-axis vibration. 4) Integrate the four basic models in step 3) using the weighted average method to obtain an integrated model; The weighting method is to weight the output class probability of the base classifier and assign the weight using the inverse mean square method; The vector output by the AlexNet-RCBAM network is converted into the probability corresponding to each category through the Softmax function, as shown below: Where n is the number of recognition task objects; a i is the i-th element of the vector input to Softmax; p i The probability that the model identifies the data as class i; The weighted harmonic mean F β The calculation formula for the precision and recall rate of the comprehensive base classifier is as follows: Where P is the precision rate; R is the recall rate; β is the weight between P and R; F β The value represents the ability of the base classifier to accurately identify a certain type of lithology, so 1-F β The value of can represent the prediction error relative to perfect recognition. Based on the principle that a large error should be assigned a small weight and a small error should be assigned a large weight, the weight is assigned using the inverse mean square method. The calculation formula is as follows: Where, F βij W is the weighted harmonic mean of the j-th base classifier identifying the i-th class object; ij The weight assigned to the j-th base classifier when identifying the i-th class object, and satisfying After integration, the probability that the strong classifier identifies the data as the i-th category is P wi =W i1 p i1 +W i2 p i2 +…+W im p im .

2. The construction method according to claim 1, wherein: In step 2), the pre-emphasis is to filter the acoustic vibration signal with a first-order FIR high-pass filter, and the transfer function of the filter in the z domain is as follows: H(z)=1-az -1 (Formula 1) Where a is the pre-emphasis coefficient, which is usually in the range of 0.9<a<1; And / or, in step 2), the wavelet threshold denoising selects the Sym8 wavelet basis function for filtering and denoising, the decomposition layer number is 9, and the threshold quantization method processes the wavelet coefficients with a soft-hard compromise threshold function, and the formula is as follows: Where w j is the wavelet coefficient of the jth layer; w jn is the wavelet coefficient after the j-th layer threshold function processing; α∈[0,1], take α=0.5; T is the threshold, and the calculation formula is as follows: Where w1 is the wavelet coefficient of the first layer; N is the length of the signal; And / or, in step 2), the bandpass filtering process is to filter the acoustic vibration signal using a sixth-order Butterworth bandpass filter, and its z-domain transfer function is as follows: Where a(m) and b(n) are coefficient vectors; N is the order of the bandpass filter; Among them, the lower limit frequency of the band-pass filter for processing sound signals is 100Hz and the upper limit frequency is 20000Hz; the lower limit frequency of the band-pass filter for processing vibration signals is 3000Hz and the upper limit frequency is 10000Hz; And / or, in step 2), the time-frequency domain analysis uses STFT to perform time-frequency analysis: during the analysis, a Hamming window w(n) with a window length of NFFT is selected to intercept the original signal x(n), and the windowed signal is Fourier transformed. After that, the window is moved by the length of the Hop Size each time, and the two adjacent windows are overlapped by the length of the Overlap. After windowing, the Fourier transform is continued, and finally a two-dimensional matrix S(ω, τ) is obtained. The calculation formula is as follows: Since the acoustic vibration signal is a power signal, its energy should be measured by the power spectrum density. Therefore, the amplitude S(ω, τ) at the corresponding time and frequency should be converted into the power spectrum density P(ω, τ). The calculation formula is as follows: Where, f r is the frequency resolution; f s is the sampling rate; N FFT The length of the Hamming window.

3. The method according to claim 1, wherein: In step 3), the data enhancement is performed in sequence using three methods: color dithering, noise perturbation, and blurring.

4. The construction method according to claim 1, wherein: In step 3), the training and tuning are based on the Tensorflow architecture; the tuning is to tune the hyperparameters, and the adjustment objects include weight initialization method, learning rate decay strategy, learning rate decay factor, learning rate and number of iterations.

5. A real-time prediction lithology model based on the acoustic vibration signal characteristics of the drilling and rock crushing process constructed by the method of any one of claims 1 to 4.

6. A system for real-time prediction of rock properties based on acoustic vibration signal characteristics during drilling and rock fragmentation, characterized in that: It includes: An input module is used to input a sound signal and a vibration signal, wherein the vibration signal includes an x-axis vibration signal, a y-axis vibration signal, and a z-axis vibration signal; A calculation module, which calculates the final prediction result of the rock mass characteristics according to the model module constructed by the method of any one of claims 1 to 4 and the information of the input module; Output module, used to output the final rock mass characteristic prediction results.

7. A method for real-time prediction of rock properties based on acoustic vibration signal characteristics during drilling and rock fragmentation, characterized by: Here are the steps: (1) obtaining a sound signal and a vibration signal of the drilled rock to be inspected, wherein the vibration signal includes an x-axis vibration signal, a y-axis vibration signal, and a z-axis vibration signal; (2) preprocessing the aforementioned sound signal and vibration signal respectively, wherein the preprocessing includes pre-emphasis, wavelet threshold noise reduction, bandpass filtering and time-frequency domain analysis to obtain time-frequency images of the sound signal, x-axis vibration signal, y-axis vibration signal and z-axis vibration signal; (3) Input the time-frequency image obtained in step 2) into the model constructed by any one of the methods of claims 1 to 4, and draw a conclusion by comparison.

8. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and the computer program is used to implement the method for real-time prediction of lithology based on acoustic vibration signal characteristics during drilling and rock crushing as described in claim 7.