Method for identifying rock CT image crack based on explicit visual prompt

By introducing explicit visual cues into deep neural networks and using high-frequency features to explicit cues, the problems of unstable accuracy of crack recognition and insufficient generalization ability in rock CT images are solved, and higher recognition accuracy and robustness are achieved.

CN120182189APending Publication Date: 2025-06-20CHINA COAL RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510228933.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The prior art uses deep neural network to identify cracks in rock CT images, and there are problems with unstable recognition accuracy and insufficient generalization capabilities, especially in noise and complex texture environments.

Method used

Using an explicit visual cues method, preprocessing the rock CT images, high-frequency features are extracted, and the explicit visual cues generator is used to embed these features as explicit prompt information into the deep neural network, guiding the network to focus on key details of the cracks.

Benefits of technology

It significantly improves the recognition accuracy of cracks in rock CT images, enhances the robustness of deep neural networks in complex texture and noise environments, and has a high degree of versatility and adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182189A_ABST
    Figure CN120182189A_ABST
Patent Text Reader

Abstract

The invention relates to the field of image data processing, in particular to a method for identifying rock CT image cracks based on explicit visual cue, which comprises the following steps: acquiring an input image; independently calculating a mask for each channel into which the image can be input; independently carrying out Fourier transform and centralization on each channel capable of inputting the image; carrying out Hadamard product on the final mask and the result after centralization; inverse centralization and inverse Fourier transform are carried out, and only the real part of the obtained result is reserved; determining a deep neural network, wherein an encoder comprises N encoding stages which are connected in sequence; constructing an explicit visual prompt generator, wherein the explicit visual prompt generator comprises N convolution modules and N pieces of prompt information which are connected in a staggered manner; and identifying the rock CT image crack. The explicit visual prompt mode not only overcomes the influence of data imbalance and noise interference on the deep neural network, but also has high universality, and can maintain good adaptability and stability in different tasks and scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image data processing, and particularly to a method for identifying fractures in rock CT images based on explicit visual cues. Background Art

[0002] Computed Tomography (CT) is a non-destructive detection technology based on X-ray imaging, which has been widely used in the field of geological engineering in recent years to obtain high-resolution images of the internal structure of rocks. By analyzing rock CT images, the physical properties of rocks, such as porosity, density, fracture distribution, etc., can be evaluated, which is of great significance for the research in the field of rock mechanics.

[0003] Computer Vision (CV) is an important branch of artificial intelligence and computer science, and its core task is to convert the original pixel data of images or videos into an understanding of high-level concepts. With the continuous development of deep learning technology, using deep neural networks to automatically extract features in images to complete various downstream tasks of computer vision has become an effective method.

[0004] However, due to the presence of noise and complex texture features in rock CT images, accurately identifying and segmenting fractures using deep neural networks remains a challenging task. Currently, mainstream deep neural network models and their improved models in the field of computer vision, such as CNNs, GANs, Transformers, etc., often show problems of unstable recognition accuracy and insufficient generalization ability due to data imbalance and noise interference in the task of identifying fractures in rock CT images. Some recent studies have tried to improve the performance of deep neural networks by implicitly adding cue information, such as attaching annotations to specific rock CT images. However, this implicit cue is difficult to control, easily leads to overfitting of the deep neural network to the cue information, and the effect of the implicit cue often depends on the deep neural network structure and is difficult to be universal among different tasks. Summary of the Invention

[0005] To solve the above problems, the present invention proposes a method for identifying fractures in rock CT images based on explicit visual cues, including the following steps:

[0006] S1: Conduct CT scanning on the rock sample and perform three-dimensional core reconstruction to obtain a three-dimensional digital core. Perform two-dimensional slicing on the three-dimensional digital core and label the fractures in each slice; preprocess to obtain an input image that can be directly input into the deep neural network;

[0007] S2: Calculate masks separately for each channel of the input image to obtain the final mask;

[0008] S3: Perform Fourier transform separately for each channel of the input image, and then centralize the transformation results.

[0009] S4: Perform Hadamard product on the final mask obtained in step S2 and the centralized matrix obtained in step S3, and denote the result as G(u, v); perform inverse centralization on G(u, v) to obtain G c (u, v).

[0010] S5: Perform inverse Fourier transform on G c (u, v) and only keep the real part of the obtained result.

[0011] S6: Determine a deep neural network, which includes an encoder and a decoder, and the encoder contains N consecutively connected encoding stages.

[0012] S7: Construct an explicit visual cue generator, which contains N interleaved convolutional modules and N pieces of cue information.

[0013] S8: Identify the fractures in the rock CT image. Add the input image and the real part result as the input to convolutional module 1 in the explicit visual cue generator to generate cue information 1, and then sequentially input cue information n to convolutional module n + 1 to generate cue information n + 1 until cue information N is generated.

[0014] Add the input image and cue information 1, and input the added result into encoding stage 1. Then sequentially input the result of adding cue information n and the output of encoding stage n - 1 into encoding stage n until it is input into encoding stage N.

[0015] Preferably, in step S1, the preprocessing method is: Cut the slices with labeled fractures into the same size and convert them into tensors. Use torchvision.transforms.Compose([]) in PyTorch to define a data preprocessing pipeline, and its parameter list contains multiple processing operations. Among them, torchvision.transforms.Resize(H, W) can cut the last two dimensions of the slices with labeled fractures into H×W, and torchvision.transforms.ToTensor() can convert the slices with labeled fractures into tensors; the input image is an RGB image, and the number of channels C = 3.

[0016] Preferably, in step S1, the input image is represented as a tensor with a dimension of C×H×W, where C is the number of channels of the image, H is the height of the image, and W is the width of the image.

[0017] Preferably, in step S2, the calculation process is: Calculate the effective covering side length L of the mask. R is the filtering rate, and its value range is a floating-point number from 0 to 1. A square with side length L is made with the geometric center O of the mask as the center. The values of the elements in the mask covered by the square are set to 0, and the values of the remaining uncovered elements are set to 1 to obtain the final mask.

[0018] Preferably, in step S2, first use torch.ones() to generate a tensor with all element values equal to 1 and the same dimension as the input image, and then set the part covered by the square to 0 through the slicing operation of the tensor.

[0019] Preferably, in step S3, each channel of the input image is a two-dimensional matrix f(x, y) of size H×W, and the Fourier transform formula is:

[0020]

[0021] where F(u, v) is a complex-valued matrix in the frequency domain, u and v respectively represent the horizontal and vertical frequencies in the frequency domain, x and y respectively represent the horizontal and vertical coordinates in the spatial domain, and j is the imaginary unit, satisfying j 2 = -1;

[0022] The centering calculation formula is: F c (u, v) = F(u, v)·(-1) u+v , where F c (u, v) is the centered matrix.

[0023] Preferably, in step S3, use torch.fft.fft2() to perform Fourier transform on each channel, and then use torch.fft.fftshift() to center the transformation result.

[0024] Preferably, in step S4, the inverse centering calculation formula is: G c (u, v) = G(u, v)·(-1) u+v , where G c (u, v) is the inverse centered matrix.

[0025] Preferably, in step S5, the calculation formula of the inverse Fourier transform is:

[0026]

[0027] Preferably, in step S5, the inverse centering is achieved by centering again, that is, still using torch.fft.fftshift() to achieve, and the inverse Fourier transform is achieved by torch.fft.ifft2().

[0028] Preferably, in step S6, each encoding stage includes a downsampling module at the front and a feature fusion module at the back.

[0029] Preferably, in step S6, the N convolutional modules are respectively convolutional module 1 to convolutional module N, the N pieces of prompt information are respectively prompt information 1 to prompt information N, the nth prompt information is after the nth convolutional module, the convolutional modules and the prompt information appear in groups, and the convolutional modules are at the front. Each convolutional module includes a convolutional layer at the front and an activation layer at the back.

[0030] Preferably, in step S7, the convolutional layers and activation layers in each convolutional module define parameters through torch.nn.Conv2d() and torch.nn.ReLU() respectively.

[0031] Beneficial effects: 1. The present invention introduces an explicit visual prompt generator, which enhances the robustness of the deep neural network in complex texture and noise environments, and thus significantly improves the recognition accuracy of fissures in rock CT images.

[0032] 2. The present invention extracts the high-frequency features of the slices through steps S1 - S5, and the explicit visual prompt generator embeds these high-frequency features as explicit prompt information into the deep neural network, thereby guiding the deep neural network to focus on the key details of the fissures.

[0033] 3. The explicit visual prompt method of the present invention not only overcomes the influence of data imbalance and noise interference on the deep neural network, but also has high generality, and can maintain good adaptability and stability in different tasks and scenarios. Description of the Drawings

[0034] Figure 1 It is a schematic structural diagram of the method for identifying fissures in rock CT images based on explicit visual prompts of the present invention; Detailed Embodiments

[0035] In the detailed embodiments section, the technical solutions of the present invention will be described in detail in conjunction with the drawings.

[0036] As Figure 1 shown, a method for identifying fissures in rock CT images based on explicit visual prompts of the present invention includes the following steps:

[0037] S1: Conduct CT scans on rock samples and perform three-dimensional core reconstruction to obtain three-dimensional digital cores. Slice the three-dimensional digital cores into two dimensions and label the fractures in each slice. Further perform preprocessing to obtain rock CT images that can be directly input into a deep neural network, defined as inputtable images. In a deep neural network written using a deep learning framework such as PyTorch, the inputtable image is represented as a tensor of dimension C×H×W, where C is the number of channels of the image, H is the height of the image (i.e., the number of vertical pixels of the image), and W is the width of the image (i.e., the number of horizontal pixels of the image).

[0038] S2: Calculate masks separately for each channel of the inputtable image. The mask dimension for each channel of the inputtable image is H×W, and the method of calculating masks for each channel of the inputtable image is the same. The calculation process is as follows: First, calculate the effective masking side length L of the mask. R is the filtering rate, a floating-point number with a value range of 0 to 1. Second, take a square with side length L centered at the geometric center O of the mask, set the values of the elements in the mask covered by the square to 0, and set the values of the remaining uncovered elements to 1 to obtain the final mask.

[0039] S3: Perform Fourier transform separately for each channel of the inputtable image, and then centralize the transformation results. Each channel of the inputtable image is a two-dimensional matrix f(x, y) of size H×W. The Fourier transform formula is:

[0040]

[0041] where F(u, v) is a complex-valued matrix in the frequency domain, u and v represent the horizontal and vertical frequencies in the frequency domain respectively, x and y represent the horizontal and vertical coordinates in the spatial domain respectively, and j is the imaginary unit, satisfying j 2 = -1;

[0042] The centralization calculation formula is: F c (u, v) = F(u, v)·(-1) u+v where F c (u, v) is the centralized matrix;

[0043] S4: Perform Hadamard product on the final mask obtained in step S2 and the centralized matrix obtained in step S3, and denote the result as G(u, v); perform inverse centralization on G(u, v). The inverse centralization calculation formula is: G c (u, v) = G(u, v)·(-1) u+v where G c (u, v) is the inverse centralized matrix;

[0044] S5: For G c(u, v) is subjected to an inverse Fourier transform and only the real part of the resulting g is retained. c The real part of (u, v), and the calculation formula for the inverse Fourier transform is:

[0045]

[0046] S6: Determine a deep neural network, which includes an encoder and a decoder. The encoder contains N sequentially connected encoding stages, namely encoding stage 1 to encoding stage N. Each encoding stage includes a previous downsampling module and a subsequent feature fusion module.

[0047] S7: Construct an explicit visual cue generator, which contains N interleaved convolutional modules and N cue messages. The N convolutional modules are respectively convolutional module 1 to convolutional module N, and the N cue messages are respectively cue message 1 to cue message N. The nth cue message follows the nth convolutional module. The convolutional modules and cue messages appear in groups, with the convolutional module in front. Each convolutional module includes a previous convolutional layer and a subsequent activation layer.

[0048] In addition, the nth cue message corresponds to the nth encoding stage, that is, cue message n corresponds to encoding stage n.

[0049] S8: Identify the fractures in the rock CT image. Add the inputtable image in step S1 to the real part result obtained in step S5, and use the added result as the input of convolutional module 1 in the display visual cue generator to generate cue message 1. Then, sequentially input the nth cue message (cue message n) into the (n + 1)th convolutional module (convolutional module n + 1) to generate the (n + 1)th cue message (cue message n + 1) until cue message N is generated.

[0050] Add the inputtable image in step S1 to cue message 1, and input the added result into encoding stage 1. Then, sequentially input the result of adding the nth cue message (cue message n) to the output of the (n - 1)th encoding stage (encoding stage n - 1) into the nth encoding stage (encoding stage n) until encoding stage N is input.

[0051] In a preferred embodiment, in step S1, the CT scanning device used for CT scanning of rock samples is the nanovoxel4000 series produced by Tianjin Sanying Precision Instrument Co., Ltd.; for three-dimensional core reconstruction, the three-dimensional digital core modeling software Avizo is used; the preprocessing method is as follows: the slices with marked fractures are cut into the same size and converted into tensors. For example, in PyTorch, torchvision.transforms.Compose([]) is used to define a data preprocessing pipeline, and its parameter list contains multiple processing operations. Among them, torchvision.transforms.Resize(H,W) can cut the last two dimensions of the slices with marked fractures into H×W, and torchvision.transforms.ToTensor() can convert the slices with marked fractures into tensors; the input image can be an RGB image, and the number of channels C = 3.

[0052] In a preferred embodiment, in step S2, first use torch.ones() to generate a tensor with all element values equal to 1 and the same dimension as the input image, and then set the masked part of the square to 0 through the slicing operation of the tensor.

[0053] In a preferred embodiment, in step S3, use torch.fft.fft2() to perform Fourier transform on each channel, and then use torch.fft.fftshift() to centralize the transformation result.

[0054] In a preferred embodiment, in step S5, inverse centralization is achieved by centralizing again, that is, still using torch.fft.fftshift() to achieve, and inverse Fourier transform is achieved by torch.fft.ifft2().

[0055] In a preferred embodiment, in step S7, the convolutional layer and activation layer in each convolutional module are defined with parameters through torch.nn.Conv2d() and torch.nn.ReLU() respectively.

[0056] The above is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Within the technical scope disclosed by the present invention, any various changes or alternative solutions that can be easily thought of by those skilled in the art should be included in the protection scope of the present invention.

Claims

1. A method for identifying cracks in rock CT images based on explicit visual cues, characterized in that: include: S1: Perform CT scanning and 3D core reconstruction on the rock sample to obtain a 3D digital core, perform 2D slices on the 3D digital core and mark the fractures in each slice; Preprocessing to obtain input images that can be directly input into deep neural networks; S2: Calculate the mask separately for each channel of the input image to obtain the final mask; S3: Perform Fourier transform on each channel of the input image separately, and then centralize the transform result; S4: Perform Hadamard product on the final mask obtained in step S2 and the centralized matrix obtained in step S3, and record the result as G(u, v); decentralize G(u, v) to obtain G c (u, v); S5: G c (u, v) performs inverse Fourier transform and retains only the real part of the result; S6: Determine a deep neural network, the deep neural network includes an encoder and a decoder, the encoder includes N encoding stages connected in sequence; S7: constructing an explicit visual cue generator, wherein the explicit visual cue generator comprises N convolutional modules and N cue information that are interlaced; S8: Identify cracks in the rock CT image, add the input image and the real part result as the input of the convolution module 1 in the explicit visual prompt generator input to generate prompt information 1, and then input the prompt information n into the convolution module n+1 in sequence to generate prompt information n+1, until prompt information N is generated; The input image is added to the prompt information 1, and the result of the addition is input to the encoding stage 1. Then, the result of adding the prompt information n and the output of the encoding stage n-1 is input to the encoding stage n in sequence, until it is input to the encoding stage N.

2. The method for identifying cracks in rock CT images based on explicit visual cues according to claim 1, characterized in that: In step S1, the preprocessing method is: cut the slices of the labeled cracks into the same size and convert them into tensors. In PyTorch, torchvision.transforms.Compose([]) is used to define a data preprocessing pipeline. Its parameter list contains multiple processing operations, among which torchvision.transforms.Resize(H,W) can cut the last two dimensions of the slices of the labeled cracks into H×W, and torchvision.transforms.ToTensor() can convert the slices of the labeled cracks into tensors; the input image can be an RGB image with a channel number C=3.

3. The method for identifying cracks in rock CT images based on explicit visual cues according to claim 2, characterized in that: In step S1, the input image may be represented as a tensor of dimension C×H×W, where C is the number of channels of the image, H is the height of the image, and W is the width of the image.

4. The method for identifying cracks in rock CT images based on explicit visual cues according to claim 3, characterized in that: In step S2, the calculation process is as follows: calculate the effective masking side length L of the mask, R is the filter rate, a floating point number ranging from 0 to 1; a square with a side length of L is made with the geometric center O of the mask as the center, and the values ​​of the elements in the mask covered by the square are set to 0, and the values ​​of the other elements not covered are set to 1 to obtain the final mask.

5. The method for identifying cracks in rock CT images based on explicit visual cues according to claim 4, characterized in that: In step S3, each channel size of the input image is a two-dimensional matrix f(x, y) of H×W, and the Fourier transform formula is: Where F(u, v) is a complex value matrix in the frequency domain, u and v represent the horizontal and vertical frequencies in the frequency domain, x and y represent the horizontal and vertical coordinates in the spatial domain, and j is an imaginary unit that satisfies j 2 = -1; The centralization calculation formula is: F c (u, v) = F(u, v)·(-1) u+v , where F c (u, v) is the centered matrix.

6. The method for identifying cracks in rock CT images based on explicit visual cues according to claim 5, characterized in that: In step S4, the inverse centralization calculation formula is: G c (u, v) = G(u, v)·(-1) u+v , where G c (u, v) is the matrix after inverse centering.

7. The method for identifying cracks in rock CT images based on explicit visual cues according to claim 6, characterized in that: In step S5, the calculation formula of the inverse Fourier transform is:

8. The method for identifying cracks in rock CT images based on explicit visual cues according to claim 1 or 7, characterized in that: In step S6, each encoding stage includes a downsampling module in front and a feature fusion module in the back.

9. The method for identifying cracks in rock CT images based on explicit visual cues according to claim 8, characterized in that: In step S6, the N convolution modules are convolution module 1 to convolution module N, and the N prompt information are prompt information 1 to prompt information N, respectively. The nth convolution module is followed by the nth prompt information. The convolution modules and the prompt information appear in groups, with the convolution modules in front. Each convolution module includes a convolution layer in front and an activation layer in the back.