Rock joint and bedding segmentation method based on frequency-space learning

By employing a frequency-space learning method, combined with dynamic convolution of rock frequency and attention cross-enhancement technology, the challenges of accuracy and efficiency in rock joint and bedding segmentation were solved, enabling real-time and accurate segmentation in field surveys and improving segmentation precision and robustness.

CN120807539AActive Publication Date: 2025-10-17NORTHEASTERN UNIV CHINA +2
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511318127.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-16
Publication Date
2025-10-17
Estimated Expiration
2045-09-16

AI Technical Summary

Technical Problem

Existing rock joint and bedding segmentation methods find it difficult to achieve the optimal balance between segmentation accuracy and efficiency, especially in field environments where low image resolution makes it difficult to extract joint and fissure features. Traditional methods are time-consuming and labor-intensive, and intelligent methods still have room for improvement in segmentation accuracy and efficiency.

Method used

A frequency-space learning-based method is adopted to achieve pixel-level segmentation of rock joints and bedding through rock frequency dynamic convolution, spatial feature calibration, convolution operation, regional attention cross enhancement and EMA attention operation, enhance the perception ability of small target objects, reduce background interference and noise influence, and capture key information.

Benefits of technology

It achieves real-time and accurate segmentation of rock joints and bedding, improves segmentation accuracy and robustness, reduces the probability of missed detection of small targets, and is suitable for field surveys to provide fast and reliable segmentation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120807539A_ABST
    Figure CN120807539A_ABST
Patent Text Reader

Abstract

The invention discloses a rock joint and bedding segmentation method based on frequency-space learning, and the method comprises the steps: carrying out the dynamic convolution operation of rock frequency, the calibration operation of spatial features, the calibration operation of spatial features, the convolution operation, the region attention cross enhancement operation, the image segmentation operation, and the EMA attention operation. Pixel-level segmentation of joints and stratification of rocks is realized, and real-time and accurate segmentation results can be provided for engineers during field exploration; the method enhances the perception capability of small target objects in the image, significantly reduces the missed detection probability of small target fractures, can weaken the influence of background interference and noise on bedding features, and carries out feature extraction on bedding structures of various forms by using receptive fields of different scales, thereby effectively capturing key information of rock bedding, and improving the accuracy of rock bedding feature extraction. And the segmentation accuracy and robustness are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of rock joint and bedding segmentation, and particularly relates to a rock joint and bedding segmentation method based on frequency-space learning. BACKGROUND

[0002] Rock is the core bearing body in the field of geological engineering, and its structural characteristics, weathering degree and fragmentation degree have important influences on the stability, safety and economy of engineering. Accurate cognition and differentiation of the macro features and structure of rock are important prerequisites for rock mechanical property evaluation, engineering geological condition analysis, geological disaster early warning and construction scheme formulation and optimization. Among the many elements of rock structure, joint and bedding have more universal and important roles.

[0003] Rock joint and bedding structure identification and feature extraction methods are mainly divided into two categories: traditional methods and intelligent methods. Before the rapid development of computer technology, scan line sampling and window sampling were the more mainstream methods. Such methods often require people to manually measure joint fissures using tools such as tape measures and compasses, and to judge bedding changes and stress directions based on experience. With the increasing demand for exploration and mining efficiency, time-consuming and labor-intensive traditional methods are increasingly unable to meet people's needs. Therefore, the identification and analysis technology of rock structure gradually transitions from manual methods to intelligent methods. Although existing methods have made certain progress in reasoning speed, how to achieve an optimal balance between segmentation accuracy and efficiency is still a problem that needs to be solved. Rock joints are widely distributed, and the cracks are small, especially in real outdoor environments, and low image resolution makes it more difficult to extract joint crack features. SUMMARY

[0004] Therefore, it is necessary to propose a rock joint and bedding segmentation method based on frequency-space learning in view of the above problems.

[0005] A rock joint and bedding segmentation method based on frequency-space learning, the method comprising:

[0006] obtaining an original rock image, performing a rock frequency dynamic convolution operation, a spatial feature calibration operation and a convolution operation on the rock image to obtain a second calibration feature;

[0007] performing a convolution operation and a spatial feature calibration operation on the second calibration feature to obtain a third calibration feature;

[0008] performing a convolution operation and a regional attention cross-enhancement operation on the third calibration feature to obtain a first attention cross-enhancement feature, and performing an up-sampling operation on the first attention cross-enhancement feature to obtain a first up-sampling feature;

[0009] Performing a splicing operation on the first up-sampled feature and the third calibration feature to obtain a first merged feature;

[0010] Performing a regional attention cross enhancement operation on the first merged feature to obtain a second attention cross enhancement feature, and performing an upsampling operation on the second attention cross enhancement feature to obtain a second upsampled feature;

[0011] Performing a splicing operation on the second upsampled feature and the second calibration feature to obtain a second merged feature;

[0012] Performing a regional attention cross enhancement operation on the second merged feature to obtain a third attention cross enhancement feature, and performing an image segmentation operation on the third attention cross enhancement feature to obtain a shallow segmentation feature;

[0013] Performing a convolution operation on the third attention cross enhancement to obtain a fourth convolution feature, and concatenating the second attention cross enhancement and the fourth convolution feature to obtain a third merged feature;

[0014] Performing a regional attention cross enhancement operation on the third merged feature to obtain a fourth attention cross enhancement feature, and performing an image segmentation operation on the fourth attention cross enhancement feature to obtain a middle-level segmentation feature;

[0015] Performing a convolution operation on the fourth attention cross enhancement to obtain a fifth convolution feature, and concatenating the fifth convolution feature and the first attention cross enhancement to obtain a fourth merged feature;

[0016] Performing a cross-stage convolution operation, an EMA attention operation, and an image segmentation operation on the fourth merged feature to obtain a deep segmentation feature;

[0017] The shallow layer segmentation features, the middle layer segmentation features and the deep layer segmentation features are combined to obtain a segmented rock feature map, and the segmented rock feature map represents the joint and bedding segmentation results of the rock.

[0018] In one embodiment, performing a rock frequency dynamic convolution operation, a spatial feature calibration operation, and a convolution operation on the rock image to obtain a second calibration feature includes:

[0019] Performing two rock frequency dynamic convolution operations on the rock image to obtain a second rock frequency dynamic feature;

[0020] performing a spatial feature calibration operation on the second rock frequency dynamic feature to obtain a first calibration feature;

[0021] performing a convolution operation on the first calibration feature to obtain a first convolution feature;

[0022] The first convolution feature is subjected to a spatial feature calibration operation to obtain a second calibration feature.

[0023] In one embodiment,

[0024] The convolution operation and the spatial feature calibration operation on the second calibration feature to obtain a third calibration feature include:

[0025] The second calibration feature is subjected to a convolution operation to obtain a second convolution feature;

[0026] The second convolution feature is subjected to a spatial feature calibration operation to obtain a third calibration feature;

[0027] The convolution operation and the regional attention cross-enhancement operation on the third calibration feature to obtain a first attention cross-enhancement feature include:

[0028] The third calibration feature is subjected to a convolution operation to obtain a third convolution feature;

[0029] The third convolution feature is subjected to a regional attention cross-enhancement operation to obtain a first attention cross-enhancement feature;

[0030] The cross-stage convolution operation, the EMA attention operation, and the image segmentation operation on the fourth merged feature to obtain a deep segmentation feature include:

[0031] The fourth merged feature is subjected to a cross-stage convolution operation to obtain a cross-stage convolution feature;

[0032] The cross-stage convolution feature is subjected to an EMA attention operation to obtain an attention feature;

[0033] The attention feature is subjected to an image segmentation operation to obtain a deep segmentation feature.

[0034] In one embodiment, the two rock frequency dynamic convolution operations on the rock image to obtain a second rock frequency dynamic feature are implemented by the following expression:

[0035] (1)

[0036] (2)

[0037] wherein, is an original rock image, is a first rock frequency dynamic feature; is a second rock frequency dynamic feature; is a rock frequency dynamic convolution operation; is a discrete Fourier transform; is a Fourier disjoint weight operation; for direction perception operation; for nuclear space modulation.

[0038] In one embodiment, the specific expression of the direction perception operation operating on the rock image x is as follows:

[0039] (3)

[0040] (4)

[0041] (5)

[0042] (6)

[0043] (7)

[0044] wherein, is the original rock image, is the direction feature; Softmax is the soft maximum operation; represents a convolution operation with a convolution kernel size of 1x1; GAP is a global average pooling operation; is the first direction processing path feature; is the second direction processing path feature; is the third direction processing path feature; is the fourth direction processing path feature; represents a convolution operation with a convolution kernel size of 3x1; represents a convolution operation with a convolution kernel size of 1x3.

[0045] In one embodiment, the specific expression of the Fourier disjoint weight operation operating on the rock image is as follows:

[0046] (8)

[0047] (9)

[0048] wherein, represents that the frequency parameters are decoupled and evenly divided into n groups using the L2 norm of the Fourier index, the frequency components controlled by each group are independent of each other and do not overlap; iDFT represents the inverse discrete Fourier transform, which maps the frequency spectrum parameters back to the spatial domain; represents that the spatial domain tensor after transformation of each group is cropped into a kxk size block; Concat is a splicing operation, which splices all the blocks to obtain a convolution kernel with different frequency information; denotes the dynamically generated attention coefficient; AP denotes average pooling, FC denotes a fully connected layer, and Sigmoid is an activation function.

[0049] In one embodiment, the kernel space modulation operates on the rock image x with the specific expression as follows:

[0050] (10)

[0051] (11)

[0052] (12)

[0053] wherein, is a global channel branch of the kernel space modulation; is a local channel branch of the kernel space modulation; AP denotes average pooling, FC denotes a fully connected layer, and Sigmoid is an activation function. denotes a one-dimensional convolution operation.

[0054] In one embodiment, the spatial feature calibration operation on the second rock frequency dynamic feature to obtain the first calibration feature is implemented by the following expression:

[0055] (13)

[0056] (14)

[0057] (15)

[0058] (16)

[0059] (17)

[0060] wherein, is a first segmentation feature of the second rock frequency dynamic feature ; is a second segmentation feature of the second rock frequency dynamic feature ; is first channel feature information; is second channel feature information; is a first channel information weight; is a second channel information weight; is the first calibration feature; denotes a channel calibration operation; denotes a spatial calibration operation; denotes element-wise multiplication.

[0061] In one embodiment, performing a spatial feature calibration operation on the first convolution feature to obtain a second calibration feature is implemented by the following expression:

[0062] (18)

[0063] (19)

[0064] (20)

[0065] (twenty one)

[0066] (twenty two)

[0067] in, is the first convolution feature The first segmentation feature of is the first convolution feature The second segmentation feature; is the third channel feature information; is the fourth channel feature information; is the third channel information weight; is the fourth channel information weight; is the second calibration feature; Indicates channel calibration operation; Represents a spatial calibration operation; Represents element-wise multiplication.

[0068] In one embodiment, performing a spatial feature calibration operation on the second convolution feature to obtain a third calibration feature is implemented by the following expression:

[0069] (twenty three)

[0070] (twenty four)

[0071] (25)

[0072] (26)

[0073] (27)

[0074] in, is the second convolution feature The first segmentation feature of is the second convolution feature The second segmentation feature; is the fifth channel feature information; is the sixth channel feature information; is the fifth channel information weight; is the sixth channel information weight; is the third calibration feature; represents a channel calibration operation; represents a spatial calibration operation; represents element-wise multiplication.

[0075] The application realizes pixel-level segmentation of rock joints and bedding, can provide real-time and accurate segmentation results for engineers during field survey, enhances the perception ability of small target objects in the image, significantly reduces the missing detection probability of small target cracks, can weaken the influence of background interference and noise on bedding features, and extracts features of various morphological bedding structures by using different scales of receptive fields, so as to effectively capture key information of rock bedding and improve the accuracy and robustness of segmentation. BRIEF DESCRIPTION OF DRAWINGS

[0076] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.

[0077] Among them:

[0078] Figure 1 is an application environment diagram of the rock joint and bedding segmentation method based on frequency-space learning in an embodiment;

[0079] Figure 2 is a flowchart of the rock joint and bedding segmentation method based on frequency-space learning in an embodiment;

[0080] Figure 3 is a flowchart of the rock joint and bedding segmentation method based on frequency-space learning in an embodiment;

[0081] Figure 4 is a flowchart of the spatial feature calibration operation in an embodiment;

[0082] Figure 5 is a flowchart of the EMA attention operation in an embodiment;

[0083] Figure 6 is an example diagram of the NEU-Rock data set in an embodiment;

[0084] Figure 7Fig. 6 is a rock image after data enhancement in one embodiment;

[0085] Figure 8 Fig. 7 is a schematic diagram of a labeling method of joint fissure and bedding characteristics in one embodiment;

[0086] Figure 9 Fig. 8 is an example diagram of Crack-seg dataset in one embodiment;

[0087] Figure 10 Fig. 9 is a curve diagram of loss function change in one embodiment;

[0088] Figure 11 Fig. 10 is a comparison diagram of segmentation results of different algorithms in one embodiment;

[0089] Figure 12 Fig. 11 is a performance comparison diagram of different deployment methods of EMA attention operation in one embodiment;

[0090] Figure 13 Fig. 12 is a visualization result of rock breaking degree quantification in one embodiment;

[0091] Figure 14 Fig. 13 is a segmentation result of YOLO-RockSolver applied to other engineering material fissures in one embodiment;

[0092] Figure 15 Fig. 14 is a structural block diagram of a computer device in one embodiment. DETAILED DESCRIPTION

[0093] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0094] Rock is the core bearing body in the field of geotechnical engineering. Its structural characteristics, weathering degree and fragmentation degree have important influence on the stability, safety and economy of engineering. Accurate cognition and recognition of the macroscopic characteristics and structure of rock are important prerequisites for the evaluation of rock mechanical properties, analysis of engineering geological conditions, early warning of geological disasters, and formulation and optimization of construction scheme. Among the many elements of rock structure, joints and bedding have more universal and important roles. The methods of rock joint and bedding structure recognition and feature extraction are mainly divided into two categories: traditional methods and intelligent methods. Before the rapid development of computer technology, scan line sampling and window sampling were the more mainstream methods. This kind of method often needs people to manually measure the joint fissure by using a tape measure, compass and other tools, and judge the bedding change and stress direction by experience. With the increasing demand for exploration and mining efficiency, the traditional method which is time-consuming and laborious can no longer meet people's needs. Therefore, the recognition and analysis technology of rock structure gradually transits from manual method to intelligent method. Although the existing methods have made certain progress in reasoning speed, how to achieve the optimal balance between segmentation accuracy and efficiency is still a problem to be solved. The joint of rock has a wide distribution range and small crack, especially in the real environment of the field, the low image resolution makes it more difficult to extract the joint crack features.

[0095] To solve the above technical problems, the rock joint and bedding segmentation method based on frequency-space learning is provided.

[0096] Figure 1 The rock joint and bedding segmentation method based on frequency-space learning in an embodiment is applied to an environment diagram. Referring to Figure 1The frequency-space learning-based rock joint and bedding segmentation method is applied to a frequency-space learning-based rock joint and bedding segmentation system. The frequency-space learning-based rock joint and bedding segmentation system includes a terminal 110 and a server 120. The terminal 110 and the server 120 are connected through a network, and the terminal 110 can be a desktop terminal or a mobile terminal, and the mobile terminal can be at least one of a mobile phone, a tablet computer, a notebook computer, etc. The server 120 can be implemented by an independent server or a server cluster composed of multiple servers. The terminal 110 is used to obtain an original rock image, perform rock frequency dynamic convolution operation, spatial feature calibration operation and convolution operation on the rock image to obtain a second calibration feature; the server 120 is used to perform convolution operation and spatial feature calibration operation on the second calibration feature to obtain a third calibration feature; perform convolution operation, regional attention cross enhancement operation on the third calibration feature to obtain a first attention cross enhancement feature, perform upsampling operation on the first attention cross enhancement feature to obtain a first upsampling feature; perform splicing operation on the upsampling feature and the third calibration feature to obtain a first merged feature; perform regional attention cross enhancement operation on the first merged feature to obtain a second attention cross enhancement feature, perform upsampling operation on the second attention cross enhancement feature to obtain a second upsampling feature; perform splicing operation on the second upsampling feature and the second calibration feature to obtain a second merged feature; perform regional attention cross enhancement operation on the second merged feature to obtain a third attention cross enhancement feature, and perform image segmentation operation on the third attention cross enhancement feature to obtain a shallow layer segmentation feature; perform convolution operation on the third attention cross enhancement to obtain a fourth convolution feature, perform splicing operation on the second attention cross enhancement and the fourth convolution feature to obtain a third merged feature; perform regional attention cross enhancement operation on the third merged feature to obtain a fourth attention cross enhancement feature, and perform image segmentation operation on the fourth attention cross enhancement feature to obtain a middle layer segmentation feature; perform convolution operation on the fourth attention cross enhancement to obtain a fifth convolution feature, perform splicing operation on the fifth convolution feature and the first attention cross enhancement to obtain a fourth merged feature; perform cross-stage convolution operation, EMA attention operation and image segmentation operation on the fourth merged feature to obtain a deep layer segmentation feature; combine the shallow layer segmentation feature, the middle layer segmentation feature and the deep layer segmentation feature to obtain a segmented rock feature map, and the segmented rock feature map represents the joint and bedding segmentation result of the rock.

[0097] As Figure 2 and Figure 3As shown, in one embodiment, a rock joint and bedding plane segmentation method based on frequency-space learning is provided. The method can be applied to both terminals and servers, and this embodiment is exemplified by application to terminals. The rock joint and bedding plane segmentation method based on frequency-space learning specifically includes the following steps:

[0098] S1: Obtain an original rock image, and perform rock frequency dynamic convolution operation, spatial feature calibration operation and convolution operation on the rock image to obtain second calibration features;

[0099] S2: Perform convolution operation and spatial feature calibration operation on the second calibration features to obtain third calibration features;

[0100] S3: Perform convolution operation and regional attention cross enhancement operation on the third calibration features to obtain first attention cross enhancement features, and perform up-sampling operation on the first attention cross enhancement features to obtain first up-sampling features;

[0101] S4: Perform splicing operation on the first up-sampling features and the third calibration features to obtain first merged features;

[0102] S5: Perform regional attention cross enhancement operation on the first merged features to obtain second attention cross enhancement features, and perform up-sampling operation on the second attention cross enhancement features to obtain second up-sampling features;

[0103] S6: Perform splicing operation on the second up-sampling features and the second calibration features to obtain second merged features;

[0104] S7: Perform regional attention cross enhancement operation on the second merged features to obtain third attention cross enhancement features, and perform image segmentation operation on the third attention cross enhancement features to obtain shallow layer segmentation features;

[0105] S8: Perform convolution operation on the third attention cross enhancement to obtain fourth convolution features, and perform splicing operation on the second attention cross enhancement and the fourth convolution features to obtain third merged features;

[0106] S9: Perform regional attention cross enhancement operation on the third merged features to obtain fourth attention cross enhancement features, and perform image segmentation operation on the fourth attention cross enhancement features to obtain middle layer segmentation features;

[0107] S10: Perform convolution operation on the fourth attention cross enhancement to obtain fifth convolution features, and perform splicing operation on the fifth convolution features and the first attention cross enhancement to obtain fourth merged features;

[0108] S11: performing a cross-stage convolution operation, an EMA attention operation and an image segmentation operation on the fourth merged feature to obtain a deep segmentation feature;

[0109] S12: merging the shallow segmentation feature, the middle segmentation feature and the deep segmentation feature to obtain a segmented rock feature map, wherein the segmented rock feature map characterizes the joint and bedding segmentation results of the rock.

[0110] Specifically, the rock frequency dynamic convolution operation maps the convolution kernel parameters to the frequency domain for dynamic adjustment and combination, thereby adaptively generating convolution kernels with different frequency domain characteristics. The spatial feature calibration operation integrates more spatial position information into rich semantic features, thereby enhancing the representation and feature extraction of small targets; the EMA attention mechanism dynamically highlights key features and suppresses redundant information, thereby accurately segmenting the bedding structure. For steps S1-S12, refer to Figure 3 The expression of each variable is as follows:

[0111] (1)

[0112] (2)

[0113] (3)

[0114] wherein, is the original rock image; is the first rock frequency dynamic feature; is the second rock frequency dynamic feature; is the first calibration feature; is the first convolution feature; is the second calibration feature; is the second convolution feature; is the third calibration feature; is the third convolution feature;x Backbone is the first attention cross-enhancement feature; is the first merged feature; is the first up-sampling feature;x A2C2f is the second attention cross-enhancement feature; is the second up-sampling feature; is the second merged feature;x Neck1 is the third attention cross-enhancement feature;x Conv4 is the fourth convolution feature; is the third merged feature;x Neck2 is the fourth attention cross-enhancement feature;x Conv5 is the fifth convolution feature; for the fourth merged feature; for the cross-stage convolution feature;x Neck3 for the attention feature; for the shallow segmentation feature; for the middle segmentation feature; for the deep segmentation feature; for the segmented rock feature map.

[0115] Through the rock frequency dynamic convolution operation, the spatial feature calibration operation, the spatial feature calibration operation, the convolution operation, the regional attention cross-enhancement operation, the image segmentation operation and the EMA attention operation, pixel-level segmentation of joints and bedding of rocks is realized, real-time and accurate segmentation results can be provided for engineers in field survey, the perception ability for small target objects in images is enhanced, the missing detection probability of small target cracks is significantly reduced, the influence of background interference and noise on bedding features can be weakened, and features of various forms of bedding structures are extracted by using receptive fields of different scales, so that key information of rock bedding is effectively captured, and the accuracy and robustness of segmentation are improved.

[0116] In one embodiment, the second calibration feature obtained by performing the rock frequency dynamic convolution operation, the spatial feature calibration operation and the convolution operation on the rock image in step S1 includes:

[0117] S101: performing twice rock frequency dynamic convolution operation on the rock image to obtain a second rock frequency dynamic feature;

[0118] S102: performing spatial feature calibration operation on the second rock frequency dynamic feature to obtain a first calibration feature;

[0119] S103: performing convolution operation on the first calibration feature to obtain a first convolution feature;

[0120] S104: performing spatial feature calibration operation on the first convolution feature to obtain a second calibration feature.

[0121] In one embodiment,

[0122] The second convolution feature obtained by performing the convolution operation on the second calibration feature in step S2 includes:

[0123] S201: performing convolution operation on the second calibration feature to obtain a second convolution feature;

[0124] S202: performing spatial feature calibration operation on the second convolution feature to obtain a third calibration feature;

[0125] For the convolution operation and the region attention cross-enhancement operation on the third calibration feature in step S3 to obtain the first attention cross-enhancement feature, the following is included:

[0126] S301: performing a convolution operation on the third calibration feature to obtain a third convolution feature;

[0127] S302: performing a region attention cross-enhancement operation on the third convolution feature to obtain a first attention cross-enhancement feature;

[0128] For the cross-stage convolution operation, the EMA attention operation, and the image segmentation operation s on the fourth merged feature in step S11 to obtain the deep segmentation feature, the following is included:

[0129] S1101: performing a cross-stage convolution operation on the fourth merged feature to obtain a cross-stage convolution feature;

[0130] S1102: performing an EMA attention operation on the cross-stage convolution feature to obtain an attention feature;

[0131] S1103: performing an image segmentation operation on the attention feature to obtain a deep segmentation feature.

[0132] In one embodiment, for the two rock frequency dynamic convolution operations on the rock image in step S101 to obtain the second rock frequency dynamic feature, the following expression is implemented:

[0133] (4)

[0134] (5)

[0135] wherein, is the original rock image, is the first rock frequency dynamic feature; is the second rock frequency dynamic feature; is the rock frequency dynamic convolution operation; is the discrete Fourier transform; is the Fourier disjoint weight operation, which constructs and optimizes the convolution weight parameters by learning mutually non-overlapping frequency components to generate standard weights; is the direction perception operation, which adopts an autonomous evolution mechanism to dynamically adapt to complex directional features in the rock image through a multi-path convolution architecture; is the kernel space modulation, which dynamically adjusts the weight response of each independent convolution by predicting a dense modulation matrix instead of a simple sparse vector to generate spatial weights.

[0136] In one embodiment, the specific expression of the direction perception operation operating on the rock image x is as follows:

[0137] (6)

[0138] (7)

[0139] (8)

[0140] (9)

[0141] (10)

[0142] wherein, is the original rock image, is the direction feature; Softmax is a soft maximum operation; represents a convolution operation with a kernel size of 1x1; GAP is a global average pooling operation; is the first direction processing path feature; is the second direction processing path feature; is the third direction processing path feature; is the fourth direction processing path feature; represents a convolution operation with a kernel size of 3x1; represents a convolution operation with a kernel size of 1x3.

[0143] In one embodiment, the specific expression of the Fourier disjoint weight operation operating on the rock image is as follows:

[0144] (11)

[0145] (12)

[0146] wherein, represents that the frequency parameters are decoupled and evenly divided into n groups using the L2 norm of the Fourier index, the frequency components controlled by each group are independent of each other and do not overlap; iDFT represents a discrete inverse Fourier transform, mapping the frequency spectrum parameters back to the spatial domain; represents that the spatial domain tensor after transformation of each group is cropped into a block with a size of kxk; Concat is a splicing operation, splicing all blocks to obtain a convolution kernel with different frequency information; represents a dynamically generated attention coefficient; AP represents average pooling, FC represents a fully connected layer, and Sigmoid is an activation function.

[0147] In one embodiment, the kernel space modulation operates on the specific expression of the rock image x as follows:

[0148] (13)

[0149] (14)

[0150] (15)

[0151] wherein, is a global channel branch of the kernel space modulation; is a local channel branch of the kernel space modulation; AP represents average pooling, FC represents a fully connected layer, and Sigmoid is an activation function; represents a one-dimensional convolution operation.

[0152] In one embodiment, the spatial feature calibration operation obtains double mapping features containing spatial information and semantic relationships through feature calibration in the spatial and channel dimensions, while being able to efficiently and accurately transmit more shallow information to a deeper network. The extraction and retention ability of small target texture features is also further enhanced, and the segmentation effect of the model on rock details, joints and bedding is also further enhanced, as shown in Figure 4 The first calibration feature obtained by performing the spatial feature calibration operation on the second rock frequency dynamic feature in step S102 is implemented by the following expression:

[0153] (16)

[0154] (17)

[0155] (18)

[0156] (19)

[0157] (21)

[0158] wherein, is a first segmentation feature of the second rock frequency dynamic feature ; is a second segmentation feature of the second rock frequency dynamic feature ; is first channel feature information; is second channel feature information; is a first channel information weight; is a second channel information weight; is the first calibration feature; Indicates channel calibration operation; Represents a spatial calibration operation; Represents element-wise multiplication.

[0159] In one embodiment, performing the spatial feature calibration operation on the first convolution feature to obtain the second calibration feature in step S104 is implemented by the following expression:

[0160] (twenty two)

[0161] (twenty three)

[0162] (twenty four)

[0163] (25)

[0164] (26)

[0165] in, is the first convolution feature The first segmentation feature of is the first convolution feature The second segmentation feature; is the third channel feature information; is the fourth channel feature information; is the third channel information weight; is the fourth channel information weight; is the second calibration feature; Indicates channel calibration operation; Represents a spatial calibration operation; Represents element-wise multiplication.

[0166] In one embodiment, performing the spatial feature calibration operation on the second convolution feature in step S202 to obtain the third calibration feature is implemented by the following expression:

[0167] (27)

[0168] (28)

[0169] (29)

[0170] (30)

[0171] (31)

[0172] in, a first segmentation feature of the second convolutional feature; a second segmentation feature of the second convolutional feature; a fifth channel feature information; a sixth channel feature information; a fifth channel information weight; a sixth channel information weight; a third calibration feature; represents a channel calibration operation; represents a spatial calibration operation; represents an element-wise multiplication.

[0173] In one embodiment, in a rock image segmentation task, the bedding boundary often presents a fuzzy characteristic due to texture similarity, uneven lighting or weathering, making it difficult for traditional segmentation algorithms to accurately distinguish adjacent rock layers, and the segmentation effect is significantly reduced. This fuzzy characteristic is mainly because the pixel features such as texture and color in the boundary region lack significant changes, making it difficult for the algorithm to accurately distinguish clear edges. At present, the success of attention mechanism in various CV tasks provides an effective solution to this problem. Attention mechanism can focus on key areas of the image and enhance the feature expression of these areas. Therefore, YOLO-Rock Solver introduces EMA attention operation, which has the advantage of efficiently modeling important context dependencies across spatial and channel dimensions and adaptively fusing multi-scale feature information. Through the EMA attention operation, the model can enhance the response weight of the subtle features at the fuzzy boundary, while suppressing the interference of irrelevant background and noise regions, thereby effectively improving the recognition ability of low-contrast and gradual change bedding boundaries.

[0174] The network structure of the EMA attention operation is shown in Figure 5 It captures feature interactions in different dimensions through a multi-scale parallel subnetwork structure, avoiding channel information reduction. Unlike traditional coordinate attention, which only embeds position information into the feature map, EMA attention operation actively models the bidirectional dependency between spatial and channel dimensions using parallel processing mechanism. This parallel design not only effectively aggregates features, but also enhances the model's ability to model long-distance dependencies, ensuring the model's lightweight while reducing computational complexity.

[0175] ​​Rock fragmentation degree is a key indicator for evaluating rock strength, stability and permeability in geological engineering, which is of great significance to engineering safety and design optimization. Especially in the field of exploration, disaster emergency or engineering reconnaissance site, geologists often need to estimate the rock fragmentation degree in real time and quickly under the condition of lacking precise instrument assistance, in order to guide immediate decision-making and risk assessment. At present, there are mature techniques that can achieve accurate identification and quantitative analysis of rock fracture geometry parameters, which have certain advantages in accuracy. However, these methods generally rely on post-processing in the laboratory, requiring a long calculation time and professional personnel, and cannot be completed immediately in the field. Therefore, there is an urgent need for a rock fragmentation degree estimation method that can be quickly implemented and immediately fed back in the field. To meet this demand, this paper proposes a real-time estimation method based on the segmentation results of rock joint fracture images. The core idea is to use the position coordinate information and relative area ratio of the segmented fracture area to construct a fast estimation model. This method aims to abandon the fine description of micro parameters and prioritize the urgent need for efficient and immediate estimation of rock fragmentation degree in the field.

[0176] The calculation method of rock fragmentation degree is shown in formula (32).

[0177] (32)

[0178] Where, is the rock fragmentation degree; is the total pixel area of the fracture mask after image segmentation; is the total pixel area of the image; is the number of fractures.

[0179] There are two cases in the calculation of rock fragmentation degree: the first is that there is only one fracture in the rock image, and the area ratio of this fracture is 10%. The second case is that there are 100 fractures in the image, and the total area ratio is 10%. The fragmentation degrees of the two kinds of rocks are very different. If only the simple relative area ratio method is used for estimation, the estimated values are the same, but there is a big difference with the true data. Therefore, we add a fracture density correction factor, that is, . The numerator is used to count the number of connected domains, and the value is larger when the fracture is denser. The denominator is used to eliminate the influence of different image sizes caused by different shooting distances on the results.

[0180] The size and quality of the data set are important factors that determine the performance of deep learning algorithms. Given the current situation that there are few public data sets for rock joint fractures and bedding structures, and the resolution is low, we have constructed the NEU-Rock data set through long-term collection and organization, as shown in Figure 6The NEU-Rock dataset includes three classes: crack, horizontal bedding, and oblique bedding, and has 1000 high-quality rock images after data augmentation. These images are all taken by artificial field shooting, some from the Jinshitan area of Dalian City and some from the Fujijie area of Fuxin City.

[0181] To expand the size of the dataset and improve the robustness of the model, we perform data augmentation on the original rock pictures taken, mainly in the following three aspects:

[0182] Geometric transformation. We use cropping, mirror flipping, and scaling to simulate the geometric changes of images caused by the diversity of image acquisition devices in real-world environments. Although rotation is also a common way of data augmentation, we did not apply it to the NEU-Rock dataset. The reason is that the images we construct are usually the entire picture of the rock, which contains the features of multiple classes such as crack, horizontal bedding, and oblique bedding. If the rotation range is too large, it will have a major impact on the directional features of horizontal bedding, thereby interfering with model training.

[0183] Random noise. We simulate real-world situations such as sensor failure and poor shooting environment by adding salt and pepper noise and Gaussian noise to the image. In terms of the scale of the noise, we use random salt and pepper noise in the range of 0.4-0.5 and random Gaussian noise in the range of 0.1-0.2.

[0184] Gray scale transformation. To simulate the diverse changes in the light conditions of the field environment, we add different degrees of gray scale transformation to the original image. The rock images after applying different ways of data augmentation are shown in Figure 7 (a)-(g) are the original image and the images after cropping, mirror flipping, scaling, salt and pepper noise, Gaussian noise, and gray scale transformation, respectively.

[0185] The commonly used data labeling software in CV tasks is Labelme and Labelimg. Among them, Labelimg is mainly used for object detection, which labels the position and category information of the target by drawing a frame around the target area. However, in the rock image segmentation task, the development form of joints and bedding is irregular, and Labelimg cannot accurately label the specific position and shape features of the target. Therefore, we use Labelme to perform pixel-level manual labeling on rock images, as shown in Figure 8 We label the joint fissure and bedding separately to reduce the visual complexity when labeling the same image. Finally, we convert the generated JSON format file after labeling into an XML format file used for model training.

[0186] To further verify the effectiveness of the proposed method, we compare it with other advanced methods on public datasets. The Crack-seg dataset contains 4029 high-quality crack images in wall and road scenes, as shown in Fig. 1. Among them, 3717 images are used for model training. The test set and validation set are composed of the remaining 112 images and 200 images, respectively. Figure 9

[0187] We first introduce the experimental setup and model performance evaluation indicators in detail. Then, we compare the results with other advanced methods on the NEU-Rock and Crack-seg datasets and visualize the comparison. Finally, we use ablation experiments to verify the effectiveness of each module and evaluate the synergy between modules.

[0188] The experiments in this paper are completed under the Pytorch 2.4.0 framework in Windows 11, using Python 3.8 as the programming language. The computer CPU is Intel Core i9 13900K, the GPU is GeForce RTX 4090, and the CUDA version is 12.1. YOLO-Rock Solver is trained for 500 epochs using the SGD optimizer, with an initial learning rate of 1×10-2, which gradually decays to 1×10-2 during training. We also use Mosaic, Mixup, and HSV methods to enhance training, and the specific hyperparameter settings are shown in Table 1. The loss function changes during training, as shown in Fig. 2. Figure 10

[0189] Table 1 Hyperparameter settings of the proposed method on the NEU-Rock dataset

[0190]

[0191] To quantitatively evaluate the effectiveness of the proposed method and compare it with other methods, we select three widely used evaluation indicators: Precision, Recall, and the mean Average Precision (mAP). Precision measures the accuracy of the model's prediction results by calculating the proportion of true positive samples among all positive samples predicted by the model. The higher the accuracy, the fewer false positives the model predicts, and the more reliable the results. Recall measures the model's ability to capture all relevant instances by calculating the proportion of true positive samples successfully predicted by the model among all true positive samples. The calculation of Precision and Recall is shown in Equations (33) and (34).

[0192] ​​ (33)

[0193] (34)

[0194] mAP is able to reflect the overall performance of the model in precision and recall by calculating the area under the Precision-Recall curve. When calculating mAP, the Intersection over Union (IoU) threshold needs to be determined first. This threshold can be set as a single value, such as 0.50 (mAP50); or a range, such as from 0.50 to 0.95, with a step of 0.05 (mAP[50,95]). In the latter case, the mAP value corresponding to each IoU threshold in the range needs to be calculated respectively, and then the average of these results is taken. Generally, when the IoU value of a predicted box and a real box is greater than or equal to 0.5, it is considered a successful match; any match below this threshold is considered unsuccessful.

[0195] To measure the effectiveness of the proposed method, we compare it with five advanced algorithms on the NEU-Rock and Crack-seg datasets, as shown in Tables 2 and 3. The algorithms used in the experiment include YOLOv8, YOLOv9, YOLOv10, YOLOv11, and YOLOv12.

[0196] Table 2 Performance comparison of different methods on the NEU-Rock dataset

[0197]

[0198] Table 3 Performance comparison of different methods on the Crack-seg dataset

[0199]

[0200] The experimental results on the two datasets show that YOLO-Rock Solver has better performance compared with the other five advanced methods. Compared with YOLOv12, the proposed method reduces the parameter amount by 1.9M and the computational amount by 0.6G, making the model lightweight while having higher accuracy. On the NEU-Rock dataset, Mask Precision and mAP[50,95] are improved by 1.6% and 4.5%, respectively. In the public Crack-seg dataset, compared with the less optimal YOLOv12, there is also a certain degree of improvement, with Mask Precision, Recall, and mAP[50,95] reaching 82.3%, 66.7%, and 69.3%, respectively.

[0201] To further visually demonstrate the superiority of YOLO-Rock Solver, we show the segmentation results of different methods on the NEU-Rock dataset in the form of pictures (confidence of 0.75). At the same time, in order to ensure the simplicity of the image, we show the results of joints and fissures separately, and only keep the segmentation mask part of the image, as shown in Figure 11

[0202] From the segmentation results of different methods, it can be seen that YOLO-Rock Solver has better detection and segmentation effect on small target joints and fissures, and the probability of missing detection is lower. In the segmentation task of stratified structure, the edge of the mask image of the proposed method is more continuous, and the stratified structure is more complete. This superior performance is due to the synergistic effect of RFDConv, SFC and EMA Attention designed according to the characteristics of rock image itself, which can ensure the real-time performance of the algorithm while enhancing the response ability to subtle cracks and fuzzy boundaries.

[0203] In this section of the experiment, we arrange and combine the key modules used by YOLO-Rock Solver, and verify the effectiveness of the proposed method and the synergistic efficiency among multiple modules on the NEU-Rock data, as shown in Table 4.

[0204] Table 4 Ablation experiment results on NEU-Rock dataset

[0205]

[0206] The experimental results show that when the three proposed modules (RFDConv, SFC, EMA) are deployed at the same time, the model has the optimal segmentation performance, and mAP50 and mAP[50,95] reach 91.1% and 45.5% respectively. When the three modules are deployed respectively, RFDConv has the best effect, which benefits from the application of its direction perception, frequency differentiation and adaptive convolution, and the precision reaches 90.8%. It is worth noting that the application of SFC reduces the parameter quantity and computational quantity of the model. We did not perform model lightweight operation in SFC, which is mainly due to the use of SFC instead of C3k2 module in the original YOLOv12 structure, which also reduces the model complexity while having the segmentation ability of subtle texture.

[0207] Further observation of the deployment of two modules shows that the deployment strategy with RFDConv has the highest accuracy, which is 0.4% and 0.3% higher than the joint deployment of SFC and EMA respectively. This further verifies the complementarity between the modules.

[0208] ​In summary, the parameter and computational complexity of YOLO-Rock Solver are only 2.63M and 9.8G, which balances all indicators while ensuring the real-time performance of the algorithm, and proves the compatibility and effectiveness of the joint deployment of the three modules.

[0209] In this section, we will deploy RFDConv in different positions in the YOLO-Rock Solver backbone network to further analyze the function and efficiency of the module in the overall network structure. The experimental results on the NEU-Crack dataset are shown in Table 5.

[0210] Table 5 Performance comparison of RFDConv deployed in different ways in the backbone network

[0211]

[0212] The experimental results show that RFDConv performs better when applied to the shallow part of the backbone network. However, when we replace all standard convolutions in the backbone network with RFDConv, the model's accuracy is lower than the baseline (YOLOv12). The main reason is that the shallow region of the backbone network captures more low-level features such as image edges, colors, and textures. At this time, the low-frequency information in the image is more distinct from the high-frequency information. Deep network focuses more on high-level semantic features, which are mainly distributed in the low-frequency region due to their translational invariance and globality. The core function of RFDConv is to separate the low-frequency information from the high-frequency information in the image. In the deep layer of the network, the low-frequency information is more abundant. If we continue to perform frequency decomposition, it will destroy the global semantic continuity contained in the deep features, and thus affect the model performance. Therefore, we deploy RFDConv in the first and second layers of the backbone network to generate dynamic convolution weights through the cascade combination of direction perception and frequency decomposition, and improve the model's expression ability for low-level features.

[0213] Considering the real-time performance requirements of rock joint and bedding segmentation tasks, we did not deploy EMA before all segmentation heads. Attention mechanisms need to consider global information of the image, and have high computational complexity. If we simply stack modules, although it will bring improvement in model accuracy, it will cause a serious decline in inference speed. Therefore, we only use these three deployment strategies to evaluate the performance of EMA in this task. We use heat maps to intuitively analyze the degree of attention of the model to key areas when deploying the module using different strategies, as shown in Figure 12 .

[0214] From the heat maps, we can see that the performance of the EMA module improves as the deployment location goes deeper, regardless of whether the task is to segment joint fissures or bedding planes. When we use scheme (c), the area that the model focuses on is closer to the real location of the target. In addition, we find that when using schemes (a) and (b) to complete the segmentation task of rock fissures, the branches in the upper right corner of the image interfere with the model to some extent, while in scheme (c) the area that the model focuses on is basically correct. The main reason for this phenomenon lies in the difference in feature levels and semantic abstraction. As the network goes deeper, the feature maps undergo multiple nonlinear transformations, and the resolution is lower. At this time, the features focus more on the overall structure and high-level semantic information of the target rather than the original pixel-level details. Therefore, deploying EMA before the last segmentation head (scheme c) can effectively distinguish target features from background noise, thus achieving more robust and accurate segmentation results.

[0215] The YOLO-Rock Solver provides three visualizations of joint fissure information, namely the total fissure area, the number of fissures, and the degree of fragmentation. We use relative size to represent the scale of the fissure, i.e., the pixel area of the fissure rather than the actual area. When calculating the actual area of the joint fissure, it is necessary to use the simulated skeleton method, which requires a large amount of computing resources. Especially in rocks with high fragmentation, the algorithm needs to traverse each fissure, which seriously affects the inference speed and has a huge impact on the real-time performance of the algorithm. The original intention of designing the fragmentation degree quantification visualization function is to help inexperienced non-geological personnel quickly and simply estimate the development degree of the rock, so we use the pixel area relative to the global image to measure the fragmentation degree of the rock. The quantification result of the rock fragmentation degree is shown in Figure 13

[0216] From the visualization results, we can see that the proposed method can accurately calculate the number and pixel area of joint fissures. In images where different fissure scales can be distinguished by the naked eye, the YOLO-Rock Solver also gives estimation results with obvious numerical differences.

[0217] Although there are slight differences between rock fissures and material fissures in other civil engineering fields, the overall features and structures are very similar, so the algorithm can be transferred to the defect detection field of steel, concrete, PVC pipes, etc. Taking the commonly used steel and concrete in construction as an example, we use the proposed method to segment the material fissures, as shown in Figure 14

[0218] ​​The results show that YOLO-Rock Solver can accurately segment cracks in non-rock materials, further proving the generalization of the algorithm while also providing a feasible solution for material defect detection in other engineering fields.

[0219] This application realizes the pixel-level segmentation of rock joints and bedding through rock frequency dynamic convolution operation, spatial feature calibration operation, spatial feature calibration operation, convolution operation, regional attention cross-enhancement operation, image segmentation operation and EMA attention operation, which can provide engineers with real-time and accurate segmentation results during field surveys; enhances the perception ability of small target objects in the image, significantly reduces the probability of missed detection of small target cracks, can weaken the impact of background interference and noise on bedding characteristics, and uses receptive fields of different scales to extract features of various forms of bedding structures, thereby effectively capturing the key information of rock bedding and improving the accuracy and robustness of segmentation.

[0220] Figure 15 FIG1 shows an internal structure diagram of a computer device in an embodiment. The computer device can be a terminal or a server. Figure 15 As shown, the computer device includes a processor, a memory and a network interface connected via a system bus. The memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium of the computer device stores an operating system and may also store a computer program. When the computer program is executed by the processor, the processor may implement a rock joint and bedding segmentation method based on frequency-space learning. The internal memory may also store a computer program. When the computer program is executed by the processor, the processor may implement a rock joint and bedding segmentation method based on frequency-space learning. Those skilled in the art will understand that Figure 15 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0221] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0222] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0223] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.

Claims

1. A rock joint and bedding segmentation method based on frequency-space learning, characterized in that: The method comprises: Acquire an original rock image, and perform a rock frequency dynamic convolution operation, a spatial feature calibration operation, and a convolution operation on the rock image to obtain a second calibration feature; performing a convolution operation and a spatial feature calibration operation on the second calibration feature to obtain a third calibration feature; Performing a convolution operation and a regional attention cross enhancement operation on the third calibration feature to obtain a first attention cross enhancement feature, and performing an upsampling operation on the first attention cross enhancement feature to obtain a first upsampled feature; Performing a splicing operation on the first up-sampled feature and the third calibration feature to obtain a first merged feature; Performing a regional attention cross enhancement operation on the first merged feature to obtain a second attention cross enhancement feature, and performing an upsampling operation on the second attention cross enhancement feature to obtain a second upsampled feature; Performing a splicing operation on the second upsampled feature and the second calibration feature to obtain a second merged feature; Performing a regional attention cross enhancement operation on the second merged feature to obtain a third attention cross enhancement feature, and performing an image segmentation operation on the third attention cross enhancement feature to obtain a shallow segmentation feature; Performing a convolution operation on the third attention cross enhancement to obtain a fourth convolution feature, and concatenating the second attention cross enhancement and the fourth convolution feature to obtain a third merged feature; Performing a regional attention cross enhancement operation on the third merged feature to obtain a fourth attention cross enhancement feature, and performing an image segmentation operation on the fourth attention cross enhancement feature to obtain a middle-level segmentation feature; Performing a convolution operation on the fourth attention cross enhancement to obtain a fifth convolution feature, and concatenating the fifth convolution feature and the first attention cross enhancement to obtain a fourth merged feature; Performing a cross-stage convolution operation, an EMA attention operation, and an image segmentation operation on the fourth merged feature to obtain a deep segmentation feature; The shallow layer segmentation features, the middle layer segmentation features and the deep layer segmentation features are combined to obtain a segmented rock feature map, and the segmented rock feature map represents the joint and bedding segmentation results of the rock.

2. The rock joint and bedding segmentation method based on frequency-space learning according to claim 1, characterized in that: The step of performing a rock frequency dynamic convolution operation, a spatial feature calibration operation, and a convolution operation on the rock image to obtain a second calibration feature includes: Performing two rock frequency dynamic convolution operations on the rock image to obtain a second rock frequency dynamic feature; performing a spatial feature calibration operation on the second rock frequency dynamic feature to obtain a first calibration feature; performing a convolution operation on the first calibration feature to obtain a first convolution feature; A spatial feature calibration operation is performed on the first convolution feature to obtain a second calibration feature.

3. The rock joint and bedding segmentation method based on frequency-space learning according to claim 1, characterized in that: The performing a convolution operation and a spatial feature calibration operation on the second calibration feature to obtain a third calibration feature includes: performing a convolution operation on the second calibration feature to obtain a second convolution feature; Performing a spatial feature calibration operation on the second convolution feature to obtain a third calibration feature; The performing a convolution operation and a regional attention cross enhancement operation on the third calibration feature to obtain a first attention cross enhancement feature includes: performing a convolution operation on the third calibration feature to obtain a third convolution feature; Performing a regional attention cross enhancement operation on the third convolutional feature to obtain a first attention cross enhancement feature; The step of performing a cross-stage convolution operation, an EMA attention operation, and an image segmentation operation on the fourth merged feature to obtain a deep segmentation feature includes: Performing a cross-stage convolution operation on the fourth merged feature to obtain a cross-stage convolution feature; Performing EMA attention operation on the cross-stage convolutional features to obtain attention features; Performing image segmentation operation on the attention feature to obtain deep segmentation feature.

4. The rock joint and bedding segmentation method based on frequency-space learning according to claim 2, characterized in that: The second rock frequency dynamic feature obtained by performing two rock frequency dynamic convolution operations on the rock image is realized by the following expression: (1) (2) in, is the original rock image, is the first rock frequency dynamic characteristic; is the second rock frequency dynamic characteristic; Dynamic convolution operation for rock frequencies; is the discrete Fourier transform; is the Fourier disjoint weight operation; For direction perception operation; is nuclear spatial modulation.

5. The rock joint and bedding segmentation method based on frequency-space learning according to claim 4, characterized in that: The specific expression of the direction-aware operation on the rock image x is as follows: (3) (4) (5) (6) (7) in, is the original rock image, is the directional feature; Softmax is the soft maximum operation; Represents a convolution operation with a convolution kernel size of 1×1; GAP is a global average pooling operation; Processing path features for the first direction; Processing path features for the second direction; Process path features for the third direction; Processing path features for the fourth direction; Represents a convolution operation with a convolution kernel size of 3×1; Represents a convolution operation with a convolution kernel size of 1×3.

6. The rock joint and bedding segmentation method based on frequency-space learning according to claim 4, characterized in that: The specific expression of the Fourier disjoint weight operation on the rock image is as follows: (8) (9) in, It means using the L2 norm of the Fourier index to sort the frequencies from low to high, decoupling the frequency parameters and evenly dividing them into n groups. The frequency components controlled by each group are independent of each other and do not overlap. iDFT stands for inverse discrete Fourier transform, which maps the spectrum parameters back to the spatial domain. Indicates that the spatial domain tensor after each group transformation is cropped into blocks of size k×k; Concat is a splicing operation that re-splices all blocks to obtain convolution kernels with different frequency information; Represents the dynamically generated attention coefficient; AP represents average pooling, FC represents the fully connected layer, and Sigmoid is the activation function.

7. The rock joint and bedding segmentation method based on frequency-space learning according to claim 4, characterized in that: The specific expression of the kernel spatial modulation operation on the rock image x is as follows: (10) (11) (12) in, The global channel branch for kernel spatial modulation; is the local channel branch of kernel spatial modulation; AP represents average pooling, FC represents fully connected layer, and Sigmoid is the activation function; Represents a one-dimensional convolution operation.

8. The rock joint and bedding segmentation method based on frequency-space learning according to claim 2, characterized in that: The spatial characteristic calibration operation of the second rock frequency dynamic characteristic to obtain the first calibration characteristic is achieved by the following expression: (13) (14) (15) (16) (17) in, Second rock frequency dynamic characteristics The first segmentation feature of Second rock frequency dynamic characteristics The second segmentation feature; is the first channel feature information; is the second channel feature information; is the first channel information weight; is the second channel information weight; is the first calibration feature; Indicates channel calibration operation; Represents a spatial calibration operation; Represents element-wise multiplication.

9. The rock joint and bedding segmentation method based on frequency-space learning according to claim 3, characterized in that: The spatial feature calibration operation performed on the first convolution feature to obtain the second calibration feature is achieved by the following expression: (18) (19) (20) (21) (22) in, is the first convolution feature The first segmentation feature of is the first convolution feature The second segmentation feature; is the third channel feature information; is the fourth channel feature information; is the third channel information weight; is the fourth channel information weight; is the second calibration feature; Indicates channel calibration operation; Represents a spatial calibration operation; Represents element-wise multiplication.

10. The rock joint and bedding segmentation method based on frequency-space learning according to claim 3, characterized in that: The spatial feature calibration operation performed on the second convolution feature to obtain the third calibration feature is achieved by the following expression: (23) (24) (25) (26) (27) in, is the second convolution feature The first segmentation feature of is the second convolution feature The second segmentation feature; is the fifth channel feature information; is the sixth channel feature information; is the fifth channel information weight; is the sixth channel information weight; is the third calibration feature; Indicates channel calibration operation; Represents a spatial calibration operation; Represents element-wise multiplication.

Citation Information

Patent Citations

  • Rock joint segmentation method and device based on global self-attention transformation network

    CN115222947A

  • Rock pore segmentation method and system based on Refinenet network model, and storage medium

    CN119006823A

  • Fracture segmentation deep learning method based on logging imaging

    CN120259332A

  • Rock debris image segmentation method based on multi-scale feature enhancement and edge perception gating

    CN120355926A

  • High-resolution image semantic segmentation network for underwater scene design

    CN120635438A