Rock joint and bedding segmentation method based on frequency-space learning

By employing a frequency-space learning method, combined with dynamic convolution of rock frequency and attention cross-enhancement technology, the challenges of accuracy and efficiency in rock joint and bedding segmentation were solved, enabling real-time and accurate segmentation in field surveys and improving segmentation precision and robustness.

CN120807539BActive Publication Date: 2025-11-11NORTHEASTERN UNIV CHINA +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511318127.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-16
Publication Date
2025-11-11
Estimated Expiration
2045-09-16

AI Technical Summary

Technical Problem

Existing methods for segmenting rock joints and bedding layers struggle to achieve the optimal balance between segmentation accuracy and efficiency. In particular, low image resolution in field environments makes it difficult to extract joint and fracture features. Traditional methods are time-consuming and labor-intensive, while intelligent methods still have room for improvement in both accuracy and efficiency.

Method used

By employing a frequency-space learning-based approach, techniques such as dynamic convolution of rock frequencies, spatial feature calibration, convolution operations, and attention cross-enhancement are used to achieve pixel-level segmentation of rock joints and bedding, enhance the perception of small target objects, reduce background interference and noise effects, and capture key information.

Benefits of technology

It enables real-time and accurate segmentation of rock joints and bedding, improves segmentation accuracy and robustness, reduces the probability of missing small targets, and is suitable for providing immediate results in field surveys.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120807539B_ABST
    Figure CN120807539B_ABST
Patent Text Reader

Abstract

This invention discloses a rock joint and bedding segmentation method based on frequency-space learning. Through dynamic convolution operation of rock frequency, spatial feature calibration operation, convolution operation, region attention cross-enhancement operation, image segmentation operation, and EMA attention operation, pixel-level segmentation of rock joints and bedding is achieved. This method can provide engineers with real-time and accurate segmentation results during field surveys. It enhances the perception of small target objects in images, significantly reduces the probability of missing small target cracks, weakens the impact of background interference and noise on bedding features, and utilizes receptive fields of different scales to extract features of various bedding structures, thereby effectively capturing key information of rock bedding and improving the accuracy and robustness of segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of rock joint and bedding segmentation technology, and in particular to a rock joint and bedding segmentation method based on frequency-space learning. Background Technology

[0002] Rocks are the core carriers in the field of geological engineering, and their structural characteristics, weathering degree, and fragmentation degree have a significant impact on the stability, safety, and economy of engineering projects. Accurately recognizing and distinguishing the macroscopic characteristics and structure of rocks is an important prerequisite for assessing rock mechanical properties, analyzing engineering geological conditions, providing early warning of geological hazards, and formulating and optimizing construction plans. Among the many elements of rock structure, joints and bedding play a more universal and important role.

[0003] Methods for identifying and extracting features of rock joints and bedding structures are mainly divided into two categories: traditional methods and intelligent methods. Before the rapid development of computer technology, scanline sampling and window sampling were the mainstream methods. These methods often required manual measurement of joints and fractures using tools such as tape measures and compasses, relying on experience to judge bedding changes and stress directions. As people's demands for exploration and mining efficiency have increased, the time-consuming and labor-intensive traditional methods have become increasingly unable to meet the needs. Therefore, the technology for identifying and analyzing rock structures has gradually transitioned from manual methods to intelligent methods. Although existing methods have made some progress in inference speed, how to achieve the optimal balance between segmentation accuracy and efficiency remains a problem that urgently needs to be solved. Rock joints are widely distributed and fractures are small, especially in real-world field environments where low image resolution makes joint and fracture feature extraction even more difficult. Summary of the Invention

[0004] Therefore, it is necessary to propose a rock joint and bedding segmentation method based on frequency-space learning to address the above problems.

[0005] A method for segmenting rock joints and bedding based on frequency-space learning, the method comprising:

[0006] Obtain the original rock image, and perform rock frequency dynamic convolution operation, spatial feature calibration operation, and convolution operation on the rock image to obtain the second calibration feature;

[0007] The third calibration feature is obtained by performing convolution and spatial feature calibration operations on the second calibration feature;

[0008] The third calibration feature is subjected to convolution and region attention cross-enhancement operations to obtain the first attention cross-enhancement feature, and the first attention cross-enhancement feature is subjected to upsampling operation to obtain the first upsampling feature;

[0009] The first upsampled feature and the third calibration feature are concatenated to obtain the first merged feature;

[0010] The first merged feature is subjected to a region attention cross-enhancement operation to obtain a second attention cross-enhancement feature, and the second attention cross-enhancement feature is subjected to an upsampling operation to obtain a second upsampled feature;

[0011] The second upsampled feature and the second calibration feature are concatenated to obtain the second merged feature;

[0012] The second merged feature is subjected to a region attention cross-enhancement operation to obtain a third attention cross-enhancement feature, and the third attention cross-enhancement feature is subjected to an image segmentation operation to obtain a shallow segmentation feature;

[0013] The third attention cross enhancement is convolved to obtain the fourth convolution feature, and the second attention cross enhancement and the fourth convolution feature are concatenated to obtain the third merged feature;

[0014] The third merged feature is subjected to a region attention cross-enhancement operation to obtain a fourth attention cross-enhancement feature, and the fourth attention cross-enhancement feature is subjected to an image segmentation operation to obtain a mid-level segmentation feature;

[0015] The fifth convolutional feature is obtained by performing a convolution operation on the fourth attention cross enhancement, and the fifth convolutional feature and the first attention cross enhancement are concatenated to obtain the fourth merged feature.

[0016] The fourth merged feature is subjected to cross-stage convolution operation, EMA attention operation and image segmentation operation to obtain deep segmentation feature;

[0017] The shallow, middle, and deep segmentation features are combined to obtain a segmented rock feature map, which represents the joint and bedding segmentation results of the rock.

[0018] In one embodiment, the process of performing rock frequency dynamic convolution, spatial feature calibration, and convolution operations on the rock image to obtain the second calibration feature includes:

[0019] The second rock frequency dynamic feature is obtained by performing two rock frequency dynamic convolution operations on the rock image;

[0020] The first calibration feature is obtained by performing a spatial feature calibration operation on the second rock frequency dynamic feature;

[0021] Perform a convolution operation on the first calibration feature to obtain the first convolution feature;

[0022] The first convolutional feature is subjected to spatial feature calibration to obtain the second calibrated feature.

[0023] In one embodiment,

[0024] The process of performing convolution and spatial feature calibration operations on the second calibration feature to obtain the third calibration feature includes:

[0025] The second calibration feature is convolved to obtain the second convolutional feature;

[0026] Perform spatial feature calibration on the second convolutional feature to obtain the third calibrated feature;

[0027] The process of performing convolution and region attention cross-enhancement operations on the third calibration feature to obtain the first attention cross-enhancement feature includes:

[0028] The third calibration feature is obtained by performing a convolution operation on the third calibration feature;

[0029] The third convolutional feature is subjected to a region attention cross-enhancement operation to obtain the first attention cross-enhancement feature;

[0030] The process of performing cross-stage convolution, EMA attention, and image segmentation operations on the fourth merged feature to obtain deep segmentation features includes:

[0031] Perform a cross-stage convolution operation on the fourth merged feature to obtain a cross-stage convolution feature;

[0032] The attention features are obtained by performing an EMA attention operation on the cross-stage convolutional features;

[0033] The attention features are subjected to image segmentation to obtain deep segmentation features.

[0034] In one embodiment, the second rock frequency dynamic feature is obtained by performing two rock frequency dynamic convolution operations on the rock image, which is achieved by the following expression:

[0035] (1)

[0036] (2)

[0037] in, Original rock image, The first rock frequency dynamic characteristic; This represents the dynamic characteristics of the second rock frequency. For rock frequency dynamic convolution operation; It is the Discrete Fourier Transform; This is a Fourier disjoint weighting operation; For orientation-aware operations; For nuclear space modulation.

[0038] In one embodiment, the specific expression for the orientation-aware operation on the rock image x is as follows:

[0039] (3)

[0040] (4)

[0041] (5)

[0042] (6)

[0043] (7)

[0044] in, Original rock image, Directional features; Softmax is the soft maximum operation; This represents a convolution operation with a kernel size of 1×1; GAP represents a global average pooling operation. Process path features for the first direction; The second direction is used to process path features; Features of the third-party processing path; Processing path features for the fourth direction; This indicates a convolution operation with a kernel size of 3×1; This indicates a convolution operation with a kernel size of 1×3.

[0045] In one embodiment, the specific expression for the Fourier disjoint weighting operation on the rock image is as follows:

[0046] (8)

[0047] (9)

[0048] in, This means that the frequency parameters are sorted from low to high using the L2 norm of the Fourier exponent, and the frequency parameters are decoupled and divided into n groups. The frequency components controlled by each group are independent and do not overlap with each other. iDFT means Inverse Discrete Fourier Transform, which maps the spectral parameters back to the spatial domain. This means that the spatial domain tensor after each group transformation is truncated into a k×k block; Concat is a concatenation operation that re-concatenates all the blocks to obtain convolution kernels with different frequency information; The attention coefficients are dynamically generated; AP represents average pooling; FC represents a fully connected layer; and Sigmoid is the activation function.

[0049] In one embodiment, the specific expression for the operation of the kernel spatial modulation on the rock image x is as follows:

[0050] (10)

[0051] (11)

[0052] (12)

[0053] in, For nuclear space modulation, the global channel branch; For local channel branches of kernel space modulation; AP represents average pooling, FC represents a fully connected layer, and Sigmoid is the activation function; This represents a one-dimensional convolution operation.

[0054] In one embodiment, the spatial feature calibration operation on the second rock frequency dynamic characteristics to obtain the first calibration feature is achieved by the following expression:

[0055] (13)

[0056] (14)

[0057] (15)

[0058] (16)

[0059] (17)

[0060] in, For the second rock frequency dynamic characteristics The first segmentation feature; For the second rock frequency dynamic characteristics The second segmentation feature; This refers to the feature information of the first channel; This refers to the feature information of the second channel; Weighting of the first channel information; Weighting of the second channel information; This is the first calibration feature; Indicates channel calibration operation; Indicates a spatial calibration operation; This indicates element-wise multiplication.

[0061] In one embodiment, the spatial feature calibration operation performed on the first convolutional feature to obtain the second calibrated feature is achieved by the following expression:

[0062] (18)

[0063] (19)

[0064] (20)

[0065] (twenty one)

[0066] (twenty two)

[0067] in, The first convolutional feature The first segmentation feature; The first convolutional feature The second segmentation feature; This refers to the third channel feature information; This is the feature information for the fourth channel; Weighting of the third channel information; Weighting of the fourth channel information; This is the second calibration feature; Indicates channel calibration operation; Indicates a spatial calibration operation; This indicates element-wise multiplication.

[0068] In one embodiment, the spatial feature calibration operation performed on the second convolutional feature to obtain the third calibrated feature is achieved through the following expression:

[0069] (twenty three)

[0070] (twenty four)

[0071] (25)

[0072] (26)

[0073] (27)

[0074] in, For the second convolution feature The first segmentation feature; For the second convolution feature The second segmentation feature; This refers to the feature information of the fifth channel; This refers to the feature information of the sixth channel; Weighting of the fifth channel information; Weighting of the sixth channel information; This is the third calibration feature; Indicates channel calibration operation; Indicates a spatial calibration operation; This indicates element-wise multiplication.

[0075] This application achieves pixel-level segmentation of rock joints and bedding, providing engineers with real-time and accurate segmentation results during field surveys. It enhances the perception of small target objects in images, significantly reduces the probability of missing small target cracks, weakens the impact of background interference and noise on bedding features, and utilizes receptive fields of different scales to extract features from various bedding structures, thereby effectively capturing key information about rock bedding and improving the accuracy and robustness of segmentation. Attached Figure Description

[0076] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0077] in:

[0078] Figure 1 This is an application environment diagram of a frequency-space learning-based rock joint and bedding segmentation method in one embodiment;

[0079] Figure 2 This is a flowchart of a rock joint and bedding segmentation method based on frequency-space learning in one embodiment;

[0080] Figure 3 This is a flowchart of a rock joint and bedding segmentation method based on frequency-space learning in one embodiment;

[0081] Figure 4 This is a flowchart of a spatial feature calibration operation in one embodiment;

[0082] Figure 5 Here is a flowchart of the EMA attention operation in one embodiment;

[0083] Figure 6 Here is an example diagram of the NEU-Rock dataset in one embodiment;

[0084] Figure 7Here is a data-enhanced image of rocks in one embodiment;

[0085] Figure 8 This is a schematic diagram illustrating a method for labeling joint and bedding features in one embodiment;

[0086] Figure 9 Here is an example diagram of the Crack-seg dataset in one embodiment;

[0087] Figure 10 Here is a graph showing the change in the loss function in one embodiment;

[0088] Figure 11 This is a comparison chart of segmentation results from different algorithms in one embodiment;

[0089] Figure 12 This is a performance comparison chart of different deployment methods for EMA attention operations in one embodiment;

[0090] Figure 13 A visualization of the quantification of rock fragmentation in one embodiment;

[0091] Figure 14 This is a sample example of how YOLO-RockSolver was applied to the segmentation of cracks in other engineering materials.

[0092] Figure 15 This is a structural block diagram of a computer device in one embodiment. Detailed Implementation

[0093] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0094] Rocks are the core carriers in the field of geological engineering, and their structural characteristics, weathering degree, and fragmentation degree have a significant impact on the stability, safety, and economy of engineering projects. Accurately recognizing and distinguishing the macroscopic characteristics and structures of rocks is a crucial prerequisite for assessing rock mechanical properties, analyzing engineering geological conditions, providing early warnings of geological hazards, and formulating and optimizing construction plans. Among the many elements of rock structure, joints and bedding play a more universal and important role. Methods for identifying and extracting features of rock joint and bedding structures are mainly divided into two categories: traditional methods and intelligent methods. Before the rapid development of computer technology, scanline sampling and window sampling were the mainstream methods. These methods often required manual measurement of joints and fractures using tools such as tape measures and compasses, relying on experience to judge bedding changes and stress directions. With increasing demands for exploration and mining efficiency, the time-consuming and labor-intensive traditional methods are increasingly unable to meet these needs. Therefore, the technology for identifying and analyzing rock structures is gradually transitioning from manual methods to intelligent methods. Although existing methods have made some progress in inference speed, achieving the optimal balance between segmentation accuracy and efficiency remains a problem that urgently needs to be solved. The joints in rocks are widely distributed and the cracks are small. Especially in real-world environments, the low image resolution makes it more difficult to extract joint and crack features.

[0095] To address the aforementioned technical problems, this application provides a method for segmenting rock joints and bedding based on frequency-space learning.

[0096] Figure 1 This is an application environment diagram of a frequency-spatial learning-based rock joint and bedding segmentation method in one embodiment. (Refer to...) Figure 1This frequency-spatial learning-based rock joint and bedding segmentation method is applied to a frequency-spatial learning-based rock joint and bedding segmentation system. The system includes a terminal 110 and a server 120. The terminal 110 and server 120 are connected via a network. The terminal 110 can be a desktop terminal or a mobile terminal, specifically a mobile phone, tablet computer, laptop computer, or at least one of these. The server 120 can be a standalone server or a server cluster consisting of multiple servers. Terminal 110 is used to acquire the original rock image, and perform rock frequency dynamic convolution, spatial feature calibration, and convolution operations on the rock image to obtain a second calibration feature; server 120 is used to perform convolution and spatial feature calibration operations on the second calibration feature to obtain a third calibration feature; perform convolution and region attention cross-enhancement operations on the third calibration feature to obtain a first attention cross-enhancement feature, and perform upsampling on the first attention cross-enhancement feature to obtain a first upsampling feature; perform a concatenation operation on the upsampling feature and the third calibration feature to obtain a first merged feature; perform a region attention cross-enhancement operation on the first merged feature to obtain a second attention cross-enhancement feature, and perform an upsampling operation on the second attention cross-enhancement feature to obtain a second upsampling feature; perform a concatenation operation on the second upsampling feature and the second calibration feature to obtain a second merged feature; perform a region attention cross-enhancement operation on the second merged feature to obtain a third attention cross-enhancement feature. The following steps are performed: First, image segmentation is applied to the third attention cross-enhancement feature to obtain shallow segmentation features. Then, a convolution operation is performed on the third attention cross-enhancement feature to obtain a fourth convolutional feature. The second attention cross-enhancement feature and the fourth convolutional feature are then concatenated to obtain a third merged feature. Next, a region attention cross-enhancement feature is applied to the third merged feature to obtain a fourth attention cross-enhancement feature. Image segmentation is then performed on the fourth attention cross-enhancement feature to obtain a mid-level segmentation feature. Finally, a convolution operation is performed on the fourth attention cross-enhancement feature to obtain a fifth convolutional feature. The fifth convolutional feature and the first attention cross-enhancement feature are then concatenated to obtain a fourth merged feature. The fourth merged feature is then subjected to cross-stage convolution, EMA attention, and image segmentation to obtain deep segmentation features. Finally, the shallow, mid-level, and deep segmentation features are merged to obtain a segmented rock feature map, which represents the joint and bedding segmentation results of the rock.

[0097] like Figure 2 and Figure 3As shown, in one embodiment, a method for segmenting rock joints and bedding based on frequency-spatial learning is provided. This method can be applied to both terminals and servers; this embodiment illustrates its application to a terminal. The specific steps of this frequency-spatial learning-based method for segmenting rock joints and bedding include:

[0098] S1: Obtain the original rock image, and perform rock frequency dynamic convolution operation, spatial feature calibration operation and convolution operation on the rock image to obtain the second calibration feature;

[0099] S2: Perform convolution and spatial feature calibration operations on the second calibration feature to obtain the third calibration feature;

[0100] S3: Perform convolution operation and region attention cross-enhancement operation on the third calibration feature to obtain the first attention cross-enhancement feature, and perform upsampling operation on the first attention cross-enhancement feature to obtain the first upsampling feature;

[0101] S4: Perform a concatenation operation on the first upsampled feature and the third calibration feature to obtain the first merged feature;

[0102] S5: Perform a region attention cross-enhancement operation on the first merged feature to obtain a second attention cross-enhancement feature, and perform an upsampling operation on the second attention cross-enhancement feature to obtain a second upsampled feature;

[0103] S6: Perform a concatenation operation on the second upsampling feature and the second calibration feature to obtain the second merged feature;

[0104] S7: Perform a region attention cross-enhancement operation on the second merged feature to obtain a third attention cross-enhancement feature, and perform an image segmentation operation on the third attention cross-enhancement feature to obtain a shallow segmentation feature;

[0105] S8: Perform a convolution operation on the third attention cross enhancement to obtain a fourth convolution feature, and concatenate the second attention cross enhancement and the fourth convolution feature to obtain a third merged feature;

[0106] S9: Perform a region attention cross-enhancement operation on the third merged feature to obtain a fourth attention cross-enhancement feature, and perform an image segmentation operation on the fourth attention cross-enhancement feature to obtain a mid-level segmentation feature;

[0107] S10: Perform a convolution operation on the fourth attention cross enhancement to obtain a fifth convolution feature, and perform a concatenation operation on the fifth convolution feature and the first attention cross enhancement to obtain a fourth merged feature;

[0108] S11: Perform cross-stage convolution operation, EMA attention operation and image segmentation operation on the fourth merged feature to obtain deep segmentation feature;

[0109] S12: The shallow layer segmentation features, the middle layer segmentation features, and the deep layer segmentation features are combined to obtain a segmented rock feature map. Characterizes the joint and bedding segmentation results of the rock.

[0110] Specifically, the rock frequency dynamic convolution operation maps convolution kernel parameters to the frequency domain for dynamic adjustment and combination, thereby adaptively generating convolution kernels with different frequency domain characteristics. The spatial feature calibration operation integrates more spatial location information into rich semantic features, thereby enhancing the representation and feature extraction of small targets; the EMA attention mechanism, by dynamically highlighting key features and suppressing redundant information, accurately segments the layered structure. For steps S1-S12, refer to... Figure 3 The expression for each variable is as follows:

[0111] (1)

[0112] (2)

[0113] (3)

[0114] in, Original rock image; The first rock frequency dynamic characteristic; This represents the dynamic characteristics of the second rock frequency. This is the first calibration feature; This is the first convolutional feature; This is the second calibration feature; This is the second convolution feature; This is the third calibration feature; The third convolutional feature; x Backbone This is a first attention cross-enhancement feature; This is the first merging feature; x is the first upsampled feature; A2C2f This is a second attention cross-enhancement feature; This is the second upsampling feature; For the second merging feature; x Neck1 This is a third attention cross-enhancement feature; x Conv4 This is the fourth convolution feature; This is the third merging feature; x Neck2 This is a fourth attention cross-enhancement feature; x Conv5 This is the fifth convolution feature; This is the fourth merging feature; For cross-stage convolutional features; x Neck3 Attention features; This is a shallow segmentation feature; This is a mid-level segmentation feature; This represents deep segmentation features; This is a segmented image of the rock features.

[0115] By employing dynamic convolution operations based on rock frequency, spatial feature calibration operations, convolution operations, region attention cross-enhancement operations, image segmentation operations, and EMA attention operations, pixel-level segmentation of rock joints and bedding is achieved. This enables engineers to obtain real-time and accurate segmentation results during field surveys. It enhances the perception of small target objects in images, significantly reduces the probability of missing small target cracks, weakens the impact of background interference and noise on bedding features, and utilizes receptive fields of different scales to extract features from various bedding structures, thereby effectively capturing key information about rock bedding and improving the accuracy and robustness of segmentation.

[0116] In one embodiment, the process of performing rock frequency dynamic convolution, spatial feature calibration, and convolution operations on the rock image in step S1 to obtain the second calibration feature includes:

[0117] S101: Perform two rock frequency dynamic convolution operations on the rock image to obtain the second rock frequency dynamic feature;

[0118] S102: Perform spatial feature calibration on the second rock frequency dynamic characteristics to obtain the first calibration feature;

[0119] S103: Perform a convolution operation on the first calibration feature to obtain the first convolution feature;

[0120] S104: Perform spatial feature calibration on the first convolutional feature to obtain the second calibration feature.

[0121] In one embodiment,

[0122] The step S2, which involves performing convolution and spatial feature calibration operations on the second calibration feature to obtain the third calibration feature, includes:

[0123] S201: Perform a convolution operation on the second calibration feature to obtain the second convolution feature;

[0124] S202: Perform spatial feature calibration on the second convolutional feature to obtain the third calibration feature;

[0125] The step S3, which involves performing a convolution operation and a region attention cross-enhancement operation on the third calibration feature to obtain the first attention cross-enhancement feature, includes:

[0126] S301: Perform a convolution operation on the third calibration feature to obtain the third convolution feature;

[0127] S302: Perform a region attention cross-enhancement operation on the third convolutional feature to obtain a first attention cross-enhancement feature;

[0128] The process of performing cross-stage convolution, EMA attention, and image segmentation operations on the fourth merged feature in step S11 to obtain deep segmentation features includes:

[0129] S1101: Perform a cross-stage convolution operation on the fourth merged feature to obtain a cross-stage convolution feature;

[0130] S1102: Perform EMA attention operation on the cross-stage convolutional features to obtain attention features;

[0131] S1103: Perform image segmentation on the attention features to obtain deep segmentation features.

[0132] In one embodiment, the second rock frequency dynamic feature obtained by performing two rock frequency dynamic convolution operations on the rock image in step S101 is achieved by the following expression:

[0133] (4)

[0134] (5)

[0135] in, Original rock image, The first rock frequency dynamic characteristic; This represents the dynamic characteristics of the second rock frequency. For rock frequency dynamic convolution operation; It is the Discrete Fourier Transform; For Fourier non-overlapping weight operations, standard weights are generated by constructing and optimizing convolution weight parameters by learning non-overlapping frequency components. For orientation-aware operations, an autonomous evolution mechanism is adopted, which dynamically adapts to complex orientation features in rock images through a multi-path convolutional architecture. For kernel spatial modulation, spatial weights are generated by dynamically adjusting the weight responses of each independent convolution by predicting a dense modulation matrix rather than a simple sparse vector.

[0136] In one embodiment, the specific expression for the orientation-aware operation on the rock image x is as follows:

[0137] (6)

[0138] (7)

[0139] (8)

[0140] (9)

[0141] (10)

[0142] in, Original rock image, Directional features; Softmax is the soft maximum operation; This represents a convolution operation with a kernel size of 1×1; GAP represents a global average pooling operation. Process path features for the first direction; The second direction is used to process path features; Features of the third-party processing path; Processing path features for the fourth direction; This indicates a convolution operation with a kernel size of 3×1; This indicates a convolution operation with a kernel size of 1×3.

[0143] In one embodiment, the specific expression for the Fourier disjoint weighting operation on the rock image is as follows:

[0144] (11)

[0145] (12)

[0146] in, This means that the frequency parameters are sorted from low to high using the L2 norm of the Fourier exponent, and the frequency parameters are decoupled and divided into n groups. The frequency components controlled by each group are independent and do not overlap with each other. iDFT means Inverse Discrete Fourier Transform, which maps the spectral parameters back to the spatial domain. This means that the spatial domain tensor after each group transformation is truncated into a k×k block; Concat is a concatenation operation that re-concatenates all the blocks to obtain convolution kernels with different frequency information; The attention coefficients are dynamically generated; AP represents average pooling; FC represents a fully connected layer; and Sigmoid is the activation function.

[0147] In one embodiment, the specific expression for the operation of the kernel spatial modulation on the rock image x is as follows:

[0148] (13)

[0149] (14)

[0150] (15)

[0151] in, For nuclear space modulation, the global channel branch; For local channel branches of kernel space modulation; AP represents average pooling, FC represents a fully connected layer, and Sigmoid is the activation function; This represents a one-dimensional convolution operation.

[0152] In one embodiment, the spatial feature calibration operation obtains dual-mapped features containing spatial information and semantic relationships through feature calibration in both spatial and channel dimensions. This allows for the efficient and accurate transfer of more shallow information to deeper network layers. The ability to extract and preserve texture features of small targets further enhances the model's segmentation performance for rock detail joints and bedding, such as... Figure 4 As shown, the first calibration feature obtained by performing spatial feature calibration on the second rock frequency dynamic feature in step S102 is achieved through the following expression:

[0153] (16)

[0154] (17)

[0155] (18)

[0156] (19)

[0157] (twenty one)

[0158] in, For the second rock frequency dynamic characteristics The first segmentation feature; For the second rock frequency dynamic characteristics The second segmentation feature; This refers to the feature information of the first channel; This refers to the feature information of the second channel; Weighting of the first channel information; Weighting of the second channel information; This is the first calibration feature; Indicates channel calibration operation; Indicates a spatial calibration operation; This indicates element-wise multiplication.

[0159] In one embodiment, the spatial feature calibration operation performed on the first convolutional feature in step S104 to obtain the second calibrated feature is achieved by the following expression:

[0160] (twenty two)

[0161] (twenty three)

[0162] (twenty four)

[0163] (25)

[0164] (26)

[0165] in, The first convolutional feature The first segmentation feature; The first convolutional feature The second segmentation feature; This refers to the third channel feature information; This is the feature information for the fourth channel; Weighting of the third channel information; Weighting of the fourth channel information; This is the second calibration feature; Indicates channel calibration operation; Indicates a spatial calibration operation; This indicates element-wise multiplication.

[0166] In one embodiment, the spatial feature calibration operation performed on the second convolutional feature in step S202 to obtain the third calibrated feature is achieved by the following expression:

[0167] (27)

[0168] (28)

[0169] (29)

[0170] (30)

[0171] (31)

[0172] in, For the second convolution feature The first segmentation feature; For the second convolution feature The second segmentation feature; This refers to the feature information of the fifth channel; This refers to the feature information of the sixth channel; Weighting of the fifth channel information; Weighting of the sixth channel information; This is the third calibration feature; Indicates channel calibration operation; Indicates a spatial calibration operation; This indicates element-wise multiplication.

[0173] In one embodiment, in rock image segmentation tasks, bedding boundaries often exhibit blurred characteristics due to texture similarity, uneven lighting, or weathering, making it difficult for traditional segmentation algorithms to accurately distinguish adjacent rock layers, resulting in a significant decrease in segmentation performance. This blurring characteristic is mainly due to the lack of significant changes in pixel features such as texture and color in the boundary region, causing the algorithm to be unable to accurately distinguish clear edges. Currently, the success of attention mechanisms in various computer vision tasks provides an effective solution to this problem. Attention mechanisms can focus on key regions of an image, enhancing the feature representation of these regions. Therefore, YOLO-Rock Solver introduces the EMA attention operation, which has the advantage of efficiently modeling important contextual dependencies across spatial and channel dimensions and adaptively fusing multi-scale feature information. Through the EMA attention operation, the model can enhance the response weights to subtle features at blurred boundaries while suppressing interference from irrelevant background and noise regions, thereby effectively improving the recognition ability of low-contrast, gradient bedding boundaries.

[0174] Network structures for EMA attention operations, such as Figure 5 As shown, it captures feature interactions across different dimensions simultaneously through a multi-scale parallel sub-network structure, avoiding the reduction of channel information. Unlike traditional coordinate attention, which only embeds positional information into the feature map, EMA attention actively models the bidirectional dependency between spatial and channel dimensions using a parallel processing mechanism. This parallel design not only effectively aggregates features but also enhances the model's ability to model long-distance dependencies, ensuring lightweight design while reducing computational complexity.

[0175] Rock fragmentation is a key indicator in geological engineering for assessing rock strength, stability, and permeability, and is of great significance for engineering safety and design optimization. Especially in field exploration, disaster emergency response, or engineering reconnaissance, geologists often need to estimate rock fragmentation in real time and quickly, without the aid of sophisticated instruments, to guide immediate decision-making and risk assessment. Currently, mature technologies exist that can accurately identify and quantify the geometric parameters of rock fractures, offering advantages in accuracy. However, these methods generally rely on post-processing in indoor environments, requiring lengthy computation times and specialized personnel, making them unsuitable for real-time fieldwork. Therefore, there is an urgent need for a method for estimating rock fragmentation that can be rapidly implemented and provide immediate feedback in the field. To address this need, this paper proposes a real-time estimation method based on rock joint and fracture image segmentation results. The core idea is to utilize the location coordinates and relative area proportions of the segmented fracture regions to construct a rapid estimation model. This method aims to abandon the detailed characterization of microscopic parameters, prioritizing the urgent need for efficient and real-time estimation of rock fragmentation under field conditions.

[0176] The method for calculating the degree of rock fragmentation is shown in formula (32).

[0177] (32)

[0178] in, The degree of rock fragmentation; The total pixel area of ​​the crack mask after image segmentation; The total pixel area of ​​the image; This represents the number of cracks.

[0179] In calculating the degree of rock fragmentation, two scenarios arise: First, the rock image contains only one fracture, with this fracture accounting for 10% of the total area. Second, the image contains 100 fractures, with their combined area accounting for 10% of the total area. The degree of rock fragmentation differs significantly between these two scenarios. If only a simple relative area ratio method is used for estimation, the resulting estimates will be the same, but will differ greatly from the actual data. Therefore, we introduce a new fracture density correction factor, namely… .molecular Used to count the number of connected components; the denser the gaps, the larger the value. Denominator Used to eliminate the influence of different image sizes at different shooting distances on the results.

[0180] The size and quality of the dataset are crucial factors determining the performance of deep learning algorithms. Given the limited publicly available datasets of rock joints, fractures, and bedding structures, and their low resolution, we constructed the NEU-Rock dataset through extensive collection and organization. Figure 6As shown, the NEU-Rock dataset includes three categories: crack, horizontal bedding, and oblique bedding, and contains 1000 high-quality rock images after data augmentation. These images were all taken manually in the field, some from Jinshitan area of ​​Dalian City and some from Fujie area of ​​Fuxin City.

[0181] To expand the dataset and improve model robustness, we performed data augmentation on the original rock images, mainly in the following three aspects:

[0182] Geometric transformations. Cropping, mirroring, and scaling are used to simulate geometric changes in images caused by various factors in real-world image acquisition devices. While rotation is a common data augmentation technique, we did not apply it to the NEU-Rock dataset. This is because our constructed images are typically full-page images of rocks, containing features from multiple categories, including cracks, horizontal bedding, and diagonal bedding. Excessive rotation would significantly impact the directional features of horizontal bedding, thus interfering with model training.

[0183] Random noise. We simulated real-world scenarios such as sensor malfunction and poor shooting environments by adding salt-and-pepper noise and Gaussian noise to the images. For the noise levels, we used random salt-and-pepper noise in the range of 0.4–0.5 and random Gaussian noise in the range of 0.1–0.2.

[0184] Grayscale transformation. To simulate the diverse lighting conditions in the field, we added different degrees of grayscale transformation to the original images. Rock images after data augmentation using different methods are shown below. Figure 7 As shown. (a)–(g) are the original image and the image after cropping, mirroring, scaling, salt and pepper noise, Gaussian noise, and grayscale transformation, respectively.

[0185] Commonly used data annotation software in computer vision (CV) tasks includes Labelme and Labelimg. Labelimg is primarily used for object detection, annotating the location and category information of objects by drawing bounding boxes over their regions. However, in rock image segmentation tasks, the development of joints and bedding is irregular, making it difficult for Labelimg to accurately annotate the specific location and shape features of targets. Therefore, we use Labelme for pixel-level manual annotation of rock images, such as... Figure 8 As shown, we label joints and fissures separately to reduce the visual complexity of labeling the same image. Finally, we convert the generated JSON file into an XML file used for model training.

[0186] To further validate the effectiveness of the proposed method, we compare it with other state-of-the-art methods using a public dataset. The Crack-seg dataset includes 4029 high-quality crack images of walls and roads, such as... Figure 9 As shown. Of these, 3717 images were used for model training. The test set and validation set consisted of the remaining 112 and 200 images, respectively.

[0187] We begin by providing a detailed overview of the experimental setup and model performance evaluation metrics. Then, we compare our methods with other state-of-the-art approaches on the NEU-Rock and Crack-seg datasets, and visualize the results. Finally, we utilize ablation experiments to validate the effectiveness of each module and evaluate the synergistic effects among them.

[0188] The experiments in this paper were conducted using the PyTorch 2.4.0 framework on Windows 11, with Python 3.8 as the programming language. The computer's CPU was an Intel Core i9 13900K, the GPU was a GeForce RTX 4090, and the CUDA version was 12.1. The YOLO-Rock Solver was trained for 500 epochs using the SGD optimizer, with an initial learning rate of 1×10⁻², which gradually decayed during training, linearly changing to 1×10⁻². We also used methods such as Mosaic, Mixup, and HSV to enhance training; the specific hyperparameter settings are shown in Table 1. The changes in the loss function during training are illustrated in Table 1. Figure 10 As shown.

[0189] Table 1. Hyperparameter settings of the proposed method on the NEU-Rock dataset.

[0190]

[0191] To quantitatively evaluate the effectiveness of the proposed method and compare it with other methods, we selected three widely used evaluation metrics: Precision, Recall, and the mean Average Precision (mAP). Precision measures the accuracy of the model's predictions by calculating the proportion of truly positive samples among those predicted as positive by the model. Higher accuracy means fewer false positives and more reliable results. Recall measures the model's ability to capture all relevant instances by calculating the proportion of truly positive samples that the model successfully predicts. The calculations for Precision and Recall are shown in equations (33) and (34).

[0192] (33)

[0193] (34)

[0194] mAP, calculated by the area under the Precision-Recall curve, reflects the overall performance of a model in terms of both precision and recall. When calculating mAP, the first step is to determine the decision threshold for the Intersection over Union (IoU). This threshold can be set as a single value, such as 0.50 (mAP50); or it can be set as a range, such as from 0.50 to 0.95, increasing in increments of 0.05 (mAP[50,95]). In the latter case, the mAP value corresponding to each IoU threshold within that range needs to be calculated separately, and then these results are averaged. Generally, a predicted bounding box is considered a successful match only when its IoU with the ground truth bounding box is greater than or equal to 0.5; any match below this threshold is considered unsuccessful.

[0195] To evaluate the effectiveness of the proposed method, we compared it with five state-of-the-art algorithms on the NEU-Rock and Crack-seg datasets, as shown in Tables 2 and 3. The algorithms used in the experiments included YOLOv8, YOLOv9, YOLOv10, YOLOv11, and YOLOv12.

[0196] Table 2 Performance comparison of different methods on the NEU-Rock dataset

[0197]

[0198] Table 3 Performance comparison of different methods on the Crack-seg dataset

[0199]

[0200] Experimental results on two datasets demonstrate that YOLO-Rock Solver outperforms the other five state-of-the-art methods. Compared to YOLOv12, the proposed method reduces the number of parameters by 1.9M and the computational cost by 0.6G, achieving higher accuracy while maintaining a lightweight model. On the NEU-Rock dataset, Mask Precision and mAP[50,95] are improved by 1.6% and 4.5%, respectively. On the publicly available Crack-seg dataset, it also shows a significant improvement over the second-best performing YOLOv12, with Mask Precision, Recall, and mAP[50,95] reaching 82.3%, 66.7%, and 69.3%, respectively.

[0201] To further demonstrate the state-of-the-art performance of YOLO-Rock Solver, we present the segmentation results of different methods on the NEU-Rock dataset as images (confidence level 0.75). To ensure image simplicity, we display the joint and fracture results separately, and only retain the segmentation mask portion of the image, as shown below. Figure 11 As shown.

[0202] The segmentation results from different methods show that YOLO-Rock Solver excels in detecting and segmenting smaller joints and cracks in rock joints, with a lower false negative probability. In the task of segmenting bedding structures, the proposed method produces more continuous mask image edges and a more complete bedding structure. This superior performance is attributed to the synergistic effect of RFDConv, SFC, and EMA Attention, designed based on the characteristics of the rock image itself. This ensures real-time performance while enhancing the response to subtle cracks and blurred boundaries.

[0203] In this experiment, we arranged and combined the key modules used in YOLO-Rock Solver and verified the effectiveness of the proposed method and the collaborative efficiency among multiple modules on the NEU-Rock dataset, as shown in Table 4.

[0204] Table 4 Ablation experimental results on the NEU-Rock dataset

[0205]

[0206] Experimental results show that the model achieves optimal segmentation performance when all three proposed modules (RFDConv, SFC, and EMA) are deployed simultaneously, with mAP50 and mAP[50,95] reaching 91.1% and 45.5%, respectively. When the three modules are deployed separately, RFDConv performs best, thanks to its orientation awareness, frequency differentiation, and adaptive convolution, achieving a precision of 90.8%. Notably, the application of the SFC module reduces the number of model parameters and computational cost. We did not perform model lightweighting in SFC, mainly because SFC replaces the C3k2 module in the original YOLOv12 structure. This reduces model complexity while maintaining the ability to segment fine textures.

[0207] Further observation of the combined deployment of the two modules revealed that the deployment strategy with RFDConv achieved the highest accuracy, with RFDConv alone outperforming the combined deployment of SFC and EMA by 0.4% and 0.3%, respectively. This further validates the complementarity between the modules.

[0208] In summary, YOLO-Rock Solver has only 2.63M parameters and 9.8G of computation, respectively. While ensuring the real-time performance of the algorithm, it achieves a balanced improvement in all metrics, proving the compatibility and effectiveness of the three modules when deployed together.

[0209] In this experiment, we deployed RFDConv at different locations in the YOLO-Rock Solver backbone network to further analyze the function and efficiency of this module in the overall network structure. The experimental results on the NEU-Crack dataset are shown in Table 5.

[0210] Table 5 Performance Comparison of RFDConv Deployment Methods in Backbone Networks

[0211]

[0212] Experimental results show that RFDConv performs better when applied to the shallow layers of the backbone network. However, when all standard convolutions in the backbone network are replaced with RFDConv, the model's accuracy is lower than the baseline (YOLOv12). This is mainly because the shallow regions of the backbone network capture more low-level features such as image edges, colors, and textures, where the distinction between low-frequency and high-frequency information is more pronounced. Deeper networks focus more on high-level semantic features, which, due to their translation invariance and globality, are primarily distributed in the low-frequency region of the frequency domain. The core function of RFDConv is to separate low-frequency and high-frequency information in the image, and low-frequency information is richer in deeper layers. Further frequency decomposition would disrupt the global semantic continuity contained in the deep features, thus impacting model performance. Therefore, we deploy RFDConv in the first and second layers of the backbone network, generating dynamic convolutional weights through a cascaded combination of orientation awareness and frequency decomposition, thereby improving the model's ability to express low-level features.

[0213] Considering the real-time performance requirements of rock joint and bedding segmentation tasks, we did not deploy EMA before all segmentation heads. Attention mechanisms require consideration of global image information, resulting in high computational complexity. While stacking modules indiscriminately might improve model accuracy, it would severely degrade inference speed. Therefore, we only used these three deployment strategies to evaluate the performance of EMA in this task. We employed heatmaps to visually analyze the model's focus on key regions when deploying the module using different strategies, such as... Figure 12 As shown.

[0214] As can be seen from the heatmap, the performance of the EMA module improves with increasing deployment depth for both joint and bedding segmentation tasks. When using scheme (c), the model focuses on regions closer to the actual location of the target. Furthermore, we found that when using schemes (a) and (b) to segment rock fissures, the tree branches in the upper right corner of the image interfered with the model, while the model focused on regions that were generally correct in scheme (c). The main reason for this phenomenon lies in the difference between the feature level and the degree of semantic abstraction. As the network deepens, the feature map undergoes multiple nonlinear transformations, resulting in lower resolution. At this point, the features focus more on the overall structure and high-level semantic information of the target, rather than the original pixel-level details. Therefore, deploying the EMA before the last segmentation head (scheme c) can effectively distinguish target features from background noise, thus achieving more robust and accurate segmentation results.

[0215] YOLO-Rock Solver provides visualizations of three joint and fracture information: total fracture area, number of fractures, and degree of fragmentation. We use relative size to represent the scale of fractures, i.e., the pixel area of ​​the fracture rather than its actual area. Calculating the actual area of ​​joints and fractures requires a simulated skeleton approach, which consumes significant computational resources. Especially in rocks with a high degree of fragmentation, the algorithm needs to traverse every fracture, severely impacting inference speed and real-time performance. Our initial intention in designing the fragmentation quantification visualization function was to help inexperienced non-geologists quickly and concisely estimate the degree of rock development; therefore, we use pixel area relative to the global image to measure the degree of rock fragmentation. The quantification results of rock fragmentation are as follows: Figure 13 As shown.

[0216] The visualization results demonstrate that the proposed method can accurately calculate the number and pixel area of ​​joints and fractures. YOLO-Rock Solver also provides estimation results with significant numerical differences in images showing different fracture sizes that are discernible to the naked eye.

[0217] While rock fissures differ slightly from material fissures in other civil engineering fields, their overall characteristics and structures are remarkably similar. Therefore, this algorithm can be transferred to defect detection in steel, concrete, PVC pipes, and other materials. Taking steel and concrete, commonly used in construction, as examples, we use the proposed method to segment material fissures, such as... Figure 14 As shown. It is worth noting that these cracks are all from untrained images; we directly detected and segmented them using weights obtained from training on rock cracks.

[0218] The results show that YOLO-Rock Solver can accurately segment cracks in non-rock materials, further demonstrating the algorithm's generalization ability and providing a feasible solution for material defect detection in other engineering fields.

[0219] This application achieves pixel-level segmentation of rock joints and bedding through dynamic convolution operations based on rock frequency, spatial feature calibration operations, convolution operations, region attention cross-enhancement operations, image segmentation operations, and EMA attention operations. This enables engineers to obtain real-time and accurate segmentation results during field surveys. It enhances the perception of small target objects in images, significantly reduces the probability of missing small target cracks, weakens the impact of background interference and noise on bedding features, and utilizes receptive fields of different scales to extract features from various bedding structures, thereby effectively capturing key information about rock bedding and improving the accuracy and robustness of segmentation.

[0220] Figure 15 An internal structural diagram of a computer device in one embodiment is shown. This computer device can specifically be a terminal or a server. Figure 15 As shown, the computer device includes a processor, memory, and network interface connected via a system bus. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and may also store a computer program. When executed by the processor, this computer program enables the processor to implement a frequency-space learning-based method for segmenting rock joints and bedding. The internal memory may also store a computer program, which, when executed by the processor, enables the processor to implement a frequency-space learning-based method for segmenting rock joints and bedding. Those skilled in the art will understand that... Figure 15 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0221] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.

[0222] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0223] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A method for segmenting rock joints and bedding based on frequency-space learning, characterized in that, The method includes: Obtain the original rock image, and perform rock frequency dynamic convolution operation, spatial feature calibration operation, and convolution operation on the rock image to obtain the second calibration feature; The third calibration feature is obtained by performing convolution and spatial feature calibration operations on the second calibration feature; The third calibration feature is subjected to convolution and region attention cross-enhancement operations to obtain the first attention cross-enhancement feature, and the first attention cross-enhancement feature is subjected to upsampling operation to obtain the first upsampling feature; The first upsampled feature and the third calibration feature are concatenated to obtain the first merged feature; The first merged feature is subjected to a region attention cross-enhancement operation to obtain a second attention cross-enhancement feature, and the second attention cross-enhancement feature is subjected to an upsampling operation to obtain a second upsampled feature; The second upsampled feature and the second calibration feature are concatenated to obtain the second merged feature; The second merged feature is subjected to a region attention cross-enhancement operation to obtain a third attention cross-enhancement feature, and the third attention cross-enhancement feature is subjected to an image segmentation operation to obtain a shallow segmentation feature; The third attention cross enhancement is convolved to obtain the fourth convolution feature, and the second attention cross enhancement and the fourth convolution feature are concatenated to obtain the third merged feature; The third merged feature is subjected to a region attention cross-enhancement operation to obtain a fourth attention cross-enhancement feature, and the fourth attention cross-enhancement feature is subjected to an image segmentation operation to obtain a mid-level segmentation feature; The fifth convolutional feature is obtained by performing a convolution operation on the fourth attention cross enhancement, and the fifth convolutional feature and the first attention cross enhancement are concatenated to obtain the fourth merged feature. The fourth merged feature is subjected to cross-stage convolution operation, EMA attention operation and image segmentation operation to obtain deep segmentation feature; The shallow, middle, and deep segmentation features are combined to obtain a segmented rock feature map, which represents the joint and bedding segmentation results of the rock.

2. The rock joint and bedding segmentation method based on frequency-space learning according to claim 1, characterized in that, The second calibration feature obtained by performing rock frequency dynamic convolution operation, spatial feature calibration operation, and convolution operation on the rock image includes: The second rock frequency dynamic feature is obtained by performing two rock frequency dynamic convolution operations on the rock image; The first calibration feature is obtained by performing a spatial feature calibration operation on the second rock frequency dynamic feature; Perform a convolution operation on the first calibration feature to obtain the first convolution feature; The first convolutional feature is subjected to spatial feature calibration to obtain the second calibrated feature.

3. The rock joint and bedding segmentation method based on frequency-space learning according to claim 1, characterized in that, The process of performing convolution and spatial feature calibration operations on the second calibration feature to obtain the third calibration feature includes: The second calibration feature is convolved to obtain the second convolutional feature; Perform spatial feature calibration on the second convolutional feature to obtain the third calibrated feature; The process of performing convolution and region attention cross-enhancement operations on the third calibration feature to obtain the first attention cross-enhancement feature includes: The third calibration feature is obtained by performing a convolution operation on the third calibration feature; The third convolutional feature is subjected to a region attention cross-enhancement operation to obtain the first attention cross-enhancement feature; The process of performing cross-stage convolution, EMA attention, and image segmentation operations on the fourth merged feature to obtain deep segmentation features includes: Perform a cross-stage convolution operation on the fourth merged feature to obtain a cross-stage convolution feature; The attention features are obtained by performing an EMA attention operation on the cross-stage convolutional features; The attention features are subjected to image segmentation to obtain deep segmentation features.

4. The rock joint and bedding segmentation method based on frequency-space learning according to claim 2, characterized in that, The second dynamic rock frequency feature is obtained by performing two dynamic convolution operations on the rock image using rock frequencies, and is achieved through the following expression: (1) (2) in, Original rock image, The first rock frequency dynamic characteristic; This represents the dynamic characteristics of the second rock frequency. For rock frequency dynamic convolution operation; It is the Discrete Fourier Transform; This is a Fourier disjoint weighting operation; For orientation-aware operations; For nuclear space modulation.

5. The rock joint and bedding segmentation method based on frequency-space learning according to claim 4, characterized in that, The specific expression for the direction-aware operation on the rock image x is as follows: (3) (4) (5) (6) (7) in, Original rock image, Directional features; Softmax is the soft maximum operation; This represents a convolution operation with a kernel size of 1×1; GAP represents a global average pooling operation. Process path features for the first direction; The second direction is used to process path features; Features of the third-party processing path; Processing path features for the fourth direction; This indicates a convolution operation with a kernel size of 3×1; This indicates a convolution operation with a kernel size of 1×3.

6. The rock joint and bedding segmentation method based on frequency-space learning according to claim 4, characterized in that, The specific expression for the Fourier disjoint weighting operation on the rock image is as follows: (8) (9) in, This means that the frequency parameters are sorted from low to high using the L2 norm of the Fourier exponent, and the frequency parameters are decoupled and divided into n groups. The frequency components controlled by each group are independent and do not overlap with each other. iDFT means Inverse Discrete Fourier Transform, which maps the spectral parameters back to the spatial domain. This means that the spatial domain tensor after each group transformation is truncated into a k×k block; Concat is a concatenation operation that re-concatenates all the blocks to obtain convolution kernels with different frequency information; The attention coefficients are dynamically generated; AP represents average pooling; FC represents a fully connected layer; and Sigmoid is the activation function.

7. The rock joint and bedding segmentation method based on frequency-space learning according to claim 4, characterized in that, The specific expression for the operation of the kernel spatial modulation on the rock image x is as follows: (10) (11) (12) in, For nuclear space modulation, the global channel branch; For local channel branches of kernel space modulation; AP represents average pooling, FC represents a fully connected layer, and Sigmoid is the activation function; This represents a one-dimensional convolution operation.

8. The rock joint and bedding segmentation method based on frequency-space learning according to claim 2, characterized in that, The spatial feature calibration operation on the second rock frequency dynamic characteristics is used to obtain the first calibration feature, which is achieved by the following expression: (13) (14) (15) (16) (17) in, For the second rock frequency dynamic characteristics The first segmentation feature; For the second rock frequency dynamic characteristics The second segmentation feature; This refers to the feature information of the first channel; This refers to the feature information of the second channel; Weighting of the first channel information; Weighting of the second channel information; This is the first calibration feature; Indicates channel calibration operation; Indicates a spatial calibration operation; This indicates element-wise multiplication.

9. The rock joint and bedding segmentation method based on frequency-space learning according to claim 3, characterized in that, The second calibrated feature is obtained by performing spatial feature calibration on the first convolutional feature, which is achieved by the following expression: (18) (19) (20) (21) (22) in, The first convolutional feature The first segmentation feature; The first convolutional feature The second segmentation feature; This refers to the third channel feature information; This is the feature information for the fourth channel; Weighting of the third channel information; Weighting of the fourth channel information; This is the second calibration feature; Indicates channel calibration operation; Indicates a spatial calibration operation; This indicates element-wise multiplication.

10. The rock joint and bedding segmentation method based on frequency-space learning according to claim 3, characterized in that, The spatial feature calibration operation performed on the second convolutional feature to obtain the third calibrated feature is achieved through the following expression: (23) (24) (25) (26) (27) in, For the second convolution feature The first segmentation feature; For the second convolution feature The second segmentation feature; This refers to the feature information of the fifth channel; This refers to the feature information of the sixth channel; Weighting of the fifth channel information; Weighting of the sixth channel information; This is the third calibration feature; Indicates channel calibration operation; Indicates a spatial calibration operation; This indicates element-wise multiplication.

Citation Information

Patent Citations

  • Rock pore segmentation method and system based on Refinenet network model, and storage medium

    CN119006823A

  • Rock debris image segmentation method based on multi-scale feature enhancement and edge perception gating

    CN120355926A