CA-GAN-based rock texture synthesis method and system
By introducing CA attention mechanism and upsampling module based on content perception in rock body texture synthesis, a CA-GAN model is constructed, which solves the problems of poor diversity and insufficient sense of reality in the existing technology, and achieves high-quality and strong natural rock body texture synthesis.
Patent Information
- Application Number
- CN202510296661.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-13
AI Technical Summary
The existing rock body texture synthesis technology has the problem of poor diversity and insufficient realism in generated pictures.
Using the CA-GAN-based rock body texture synthesis method, a rock body texture synthesis model is constructed in the architecture of the generative adversarial network by introducing a CA attention mechanism and a content-aware upsampling module, and a generator and discriminator are trained to generate high-quality rock body textures.
It improves the flexibility and generalization ability of rock body texture synthesis, and the generated texture is more natural and realistic, with rich details, and is suitable for complex landform changes or body texture image generation of different rock types.
Smart Images

Figure CN120147500A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of rock texture synthesis, and particularly to a method and system for rock texture synthesis based on CA-GAN. Background Art
[0002] Three-dimensional rock models can express the mechanical properties, fracture processes, and geometric characteristics and distributions of mineral particles of rocks, and are widely used in fields such as earth science, environmental engineering, and geotechnical engineering. High-fidelity non-uniform natural rock models can accurately reproduce the morphology, size, topological relationships, and distribution laws of mineral particles. However, traditional texture mapping techniques are difficult to comprehensively represent the distribution and topological relationships of mineral particles inside and on the surface of three-dimensional models, and it is difficult to obtain natural rock textures, resulting in insufficient realism of the models.
[0003] In the prior art, rock texture synthesis methods are divided into procedural generation methods and sample generation methods. Procedural generation methods are usually limited to specific features and have problems such as insufficient diversity, overfitting, inability to adapt to complex scenarios, and loss of details. Although sample generation methods can maintain local consistency, there are still certain limitations in large-scale texture generation, global consistency, texture diversity, and the ability to adapt to complex scenarios.
[0004] Therefore, there is an urgent need to propose a method and system for rock texture synthesis based on CA-GAN, which can generate rock textures of any size and non-uniformity, so as to enhance the flexibility and generalization ability of rock texture synthesis, and generate more natural, realistic, and detailed texture maps. Summary of the Invention
[0005] In view of this, the present invention provides a method and system for rock texture synthesis based on CA-GAN to solve the technical problems of poor diversity and insufficient realism in the existing rock texture synthesis technology.
[0006] To achieve the above technical objectives, the present invention adopts the following technical solutions:
[0007] On the one hand, the present invention provides a method for rock texture synthesis based on CA-GAN, including:
[0008] Collect rock image data, preprocess the rock image data to obtain a rock texture image sample data set;
[0009] Introduce a CA attention mechanism and a content-aware upsampling module into the architecture of the generative adversarial network to construct a rock texture synthesis model based on CA-GAN;
[0010] Use the rock texture image sample data set to train the rock texture synthesis model to obtain a well-trained rock texture synthesis model;
[0011] Generate a rock texture synthesis map by training a complete rock texture synthesis model, and visually display the rock texture synthesis map.
[0012] Furthermore, introduce the CA attention mechanism and the content-aware upsampling module into the architecture of the generative adversarial network to construct an initial rock texture synthesis model based on the generative adversarial network, including:
[0013] Introduce a hybrid dilated convolution layer, a CA attention mechanism, and a content-aware upsampling module into the generator of the generative adversarial network; extract feature information of different sizes in the rock image data through the hybrid dilated convolution layer; enlarge the size of the input feature map through the content-aware upsampling module; obtain attention weights in the three directions of length, width, and height of the rock texture through the CA attention mechanism, encode the texture position information, and restore the details of the image to generate a rock texture with rich details;
[0014] Adopt discriminators of multiple different scales in the discriminator of the generative adversarial network. Each scale of discriminator is used to judge the authenticity of the image at the corresponding scale, and the judgment results of different scales are integrated to obtain a comprehensive evaluation of the generated image.
[0015] Furthermore, the structure of the generator of the rock texture synthesis model includes: a multi-scale input module, a first hybrid dilated convolution block, a second hybrid dilated convolution block, an upsampling module, and a CA-3D attention convolution module;
[0016] The input module of each scale is connected to the corresponding first hybrid dilated convolution block, and is used to input multi-scale random body noise to make the model output different textures;
[0017] The first hybrid dilated convolution block is used to process the input noise and generate temporary feature maps of different scales;
[0018] The output end of the first hybrid dilated convolution block with the smallest size is connected to the input end of the upsampling module, and the output ends of the first hybrid dilated convolution blocks of the remaining sizes are connected to the second hybrid dilated convolution block, and the output end of the second hybrid dilated convolution block is connected to the input end of the upsampling module;
[0019] The upsampling module expands the temporary feature map through a content-aware upsampling method;
[0020] Except for the temporary feature map with the smallest size, the remaining temporary feature maps of each size and the output feature maps of the upsampling module of the adjacent secondary size are used as input data and input into the second hybrid dilated convolution block, and the second hybrid dilated convolution block is used to fuse the temporary feature maps of different sizes;
[0021] The output end of the second hybrid dilated convolution block with the largest size is connected to the input end of the CA-3D attention convolution module;
[0022] The CA-3D attention convolution module is used to obtain the attention weights in the three directions of the length, width, and height of the rock body texture, transfer the entity channels to three channels, and encode the position information.
[0023] Furthermore, the first hybrid dilated convolution block and the second hybrid dilated convolution block have the same structure, both including three convolutional layers;
[0024] Among them, the convolutional kernels of the first and third convolutional layers are 1×1×1, and the second convolutional layer is a convolutional layer with a dilation rate of 2.
[0025] Furthermore, the CA-3D attention convolution module includes a pooling layer, a feature concatenation convolutional layer, a normalization activation layer, a feature decomposition layer, and an output layer connected in sequence;
[0026] The pooling layer is used to perform global average pooling on the input original feature map in the three directions of length, width, and height to obtain feature maps in the three directions;
[0027] The feature concatenation convolutional layer is used to concatenate the feature maps in the three directions of length, width, and height, and use a convolutional module with a convolutional kernel of 1×1×1 to reduce the channel dimension to obtain a low-dimensional feature map;
[0028] The normalization activation layer is used to perform a normalization operation on the low-dimensional feature map and a non-linear operation on the normalized feature map;
[0029] The feature decomposition layer is used to perform a segmentation operation along the spatial dimension to restore the feature maps in the three directions, use a 1×1×1 convolutional module to restore the channel dimension, and obtain the attention weights of the feature maps in the three directions through an activation function;
[0030] The output layer is used to weight the original feature map and the attention weights of the feature maps in the three directions, and convert the channel dimension to three dimensions through a 1×1×1 convolutional module.
[0031] Furthermore, the upsampling module includes an upsampling layer, a three-dimensional convolutional layer, a BN layer, and a Relu activation layer connected in sequence;
[0032] Among them, the upsampling layer includes a prediction unit and a feature recombination unit;
[0033] The prediction module compresses the number of channels through 1×1 convolution, converts the number of channels through content encoding, and expands the channel dimension in the spatial dimension to generate an upsampling kernel, and normalizes the upsampling kernel to keep the original feature content unchanged to obtain an output feature map;
[0034] The feature recombination module is used to map the local information of the output feature map to the input feature map, complete the recombination of the feature map, and output the upsampling output value.
[0035] Furthermore, each discriminator of each scale is used to judge the authenticity of the image at the corresponding scale, including:
[0036] Each discriminator of each scale is used to perform texture discrimination on the slice data extracted after slicing the volume data at a preset scale and the real 2D sample data;
[0037] Among them, the slice data is obtained by sampling the volume texture generated by the generator, and different scales of volume textures are formed for slicing; the input real 2D samples are adjusted to different scales through sampling operations to achieve discrimination at different scales, and the real 2D sample data and the slice data in each discriminator have the same resolution.
[0038] Furthermore, the acquisition of rock image data and the preprocessing of the rock image data to obtain a sample data set include:
[0039] Obtain real rock images with a preset resolution and screen the rock images;
[0040] Manually crop the texture areas in the screened rock images;
[0041] Resample the cropped images to generate texture pictures with different resolutions, and generate a rock texture image sample data set.
[0042] Furthermore, using the rock texture image sample data set to train the rock volume texture synthesis model to obtain a trained complete rock volume texture synthesis model, including:
[0043] Set the learning rates of the generator and the discriminator respectively, and use the AdamW optimizer for training;
[0044] A learning rate decay strategy is introduced in the training process. After each preset number of training times, the learning rate is halved to optimize the training effect of the model.
[0045] On the other hand, the present invention also provides a rock volume texture synthesis system based on CA-GAN, including:
[0046] A data acquisition module, which is used to collect rock image data and preprocess the rock image data to obtain a rock texture image sample data set;
[0047] A model construction module, which is used to introduce a CA attention mechanism and a content-aware upsampling module into the architecture of the generative adversarial network to construct a rock volume texture synthesis model based on CA-GAN;
[0048] A model training module for training the rock body texture synthesis model using a rock texture image sample dataset to obtain a completely trained rock body texture synthesis model;
[0049] An image synthesis module for generating a rock body texture synthesis map through the completely trained rock body texture synthesis model and visually displaying the rock texture synthesis map.
[0050] Compared with the prior art, the rock body texture synthesis method based on CA-GAN of the present invention has the following beneficial effects: By constructing a rock body texture synthesis model based on CA-GAN, it can adaptively adjust the key areas in the image generation process, thereby improving the fidelity of local details, effectively avoiding repeated texture patterns and over-smoothing problems, and generating more natural and diverse textures. Through the content-aware upsampling module, it is adjusted according to the content of the input rock texture sample, making the generated high-resolution texture richer and retaining the details and natural transitions in the sample, which is applicable to the generation of body texture images with complex landform changes or different rock types. The present invention can better learn the texture feature distribution on the two-dimensional texture image and extend it to the three-dimensional space, showing good consistency in different directions in the three-dimensional space, effectively solving the problems of detail loss, texture repetition, insufficient local consistency, and poor global consistency in texture generation in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 It is a schematic flowchart of the rock body texture synthesis method based on CA-GAN provided by the present invention;
[0052] Figure 2 It is a schematic structural diagram of the generator of the rock body texture synthesis model provided by the present invention;
[0053] Figure 3 It is a schematic structural diagram of the hybrid dilated convolution block provided by the present invention;
[0054] Figure 4 It is a schematic structural diagram of the CA-3D attention convolution module provided by the present invention;
[0055] Figure 5 It is a schematic structural diagram of the content-aware upsampling module provided by the present invention;
[0056] Figure 6 It is a schematic structural diagram of the discriminator of the rock body texture synthesis model provided by the present invention;
[0057] Figure 7 It is a schematic flowchart of the construction process of the rock texture image sample dataset provided by the present invention;
[0058] Figure 8 The rock texture pictures used in the verification experiment provided by the present invention;
[0059] Figure 9 The generated result diagram of the rock body texture provided by the present invention;
[0060] Figure 10 The structural schematic diagram of the rock body texture synthesis system based on CA-GAN provided by the present invention. Specific embodiments
[0061] The following combines the accompanying drawings to specifically describe the preferred embodiments of the present invention. Among them, the accompanying drawings constitute a part of this application and are used together with the embodiments of the present invention to explain the principle of the present invention, rather than to limit the scope of the present invention.
[0062] Please refer to Figure 1 , this embodiment provides a rock body texture synthesis method based on CA-GAN, including:
[0063] Step S101: Collect rock image data, preprocess the rock image data, and obtain a rock texture image sample data set;
[0064] Step S102: Introduce a CA attention mechanism and a content-aware upsampling module into the architecture of the generative adversarial network to construct a rock body texture synthesis model based on CA-GAN;
[0065] Step S103: Use the rock texture image sample data set to train the rock body texture synthesis model to obtain a trained complete rock body texture synthesis model;
[0066] Step S104: Use the trained complete rock body texture synthesis model to generate a rock body texture synthesis map and visually display the rock texture synthesis map.
[0067] The method of this embodiment can construct a rock body texture synthesis model by combining a generative adversarial network, a CA attention mechanism, and a content-aware upsampling module, and can generate high-quality and diverse rock textures while retaining details and a sense of reality. This method has significant advantages in training efficiency, synthesis effect, and application value, and is suitable for actual application scenarios.
[0068] As a preferred embodiment, the introducing a CA attention mechanism and a content-aware upsampling module into the architecture of the generative adversarial network to construct an initial rock body texture synthesis model based on the generative adversarial network includes:
[0069] Introduce a hybrid dilated convolution layer, a CA attention mechanism, and a content-aware upsampling module into the generator of the generative adversarial network; extract feature information of different sizes in the rock image data through the hybrid dilated convolution layer; enlarge the size of the input feature map through the content-aware upsampling module; obtain attention weights in the three directions of length, width, and height of the rock body texture through the CA attention mechanism, encode the texture position information, and restore the detailed information of the image to generate a rock body texture with rich details.
[0070] In the discriminator of the generative adversarial network, multiple discriminators of different scales are adopted. Each scale of the discriminator is used to judge the authenticity of the image at the corresponding scale, and the judgment results of different scales are integrated to obtain a comprehensive evaluation of the generated image.
[0071] As a preferred embodiment, the structure of the generator of the rock body texture synthesis model includes: a multi-scale input module, a first hybrid dilated convolution block, a second hybrid dilated convolution block, an upsampling module, and a CA-3D attention convolution module.
[0072] Each scale of the input module is connected to the corresponding first hybrid dilated convolution block, and is used to input multi-scale random body noise to enable the model to output different textures.
[0073] The first hybrid dilated convolution block is used to process the input noise and generate temporary feature maps of different scales.
[0074] The output end of the first hybrid dilated convolution block with the smallest size is connected to the input end of the upsampling module, and the output ends of the remaining first hybrid dilated convolution blocks of each size are connected to the second hybrid dilated convolution block, and the output end of the second hybrid dilated convolution block is connected to the input end of the upsampling module.
[0075] The upsampling module expands the temporary feature map through a content-aware upsampling method.
[0076] Except for the temporary feature map with the smallest size, the remaining temporary feature maps of each size and the output feature maps of the upsampling module of the adjacent secondary size are used as input data and input into the second hybrid dilated convolution block, and the second hybrid dilated convolution block is used to fuse the temporary feature maps of different sizes.
[0077] The output end of the second hybrid dilated convolution block with the largest size is connected to the input end of the CA-3D attention convolution module.
[0078] The CA-3D attention convolution module is used to obtain the attention weights in the three directions of length, width, and height of the rock body texture, transfer the entity channels to three channels, and encode the position information.
[0079] By combining multi-scale noise input, dilated convolution, content-aware upsampling, and a CA-3D attention convolution module, the generator can effectively improve the quality and diversity of rock texture synthesis. It can not only capture the detailed levels and global features of the texture but also avoid common problems such as texture repetition, blurring, or distortion, generating textures with high quality and a natural feel.
[0080] The following combines Figure 2 to elaborate in detail on the structure of the generator of the rock texture synthesis model.
[0081] As Figure 2 shown, in the generator of the model, a set of multi-scale random volumetric noise is input through a multi-scale input layer, so that each generated output is different to increase the diversity of the generated volumetric texture. Figure 2 In it, {c 0 , …, c K} represents the additional values required for each spatial dimension dependence, and these additional values depend on the network architecture of the generator. In the architecture adopted in this embodiment, the additional values are set to {4, 5, 6, 6, 6, 4} (taking k = 5 as an example).
[0082] The input noise is processed through the first convolution block to generate temporary feature maps of different scales. Except for the feature map with the smallest size, the feature maps of other sizes need to be further processed through the second dilated convolution block after being connected to the channels of their secondary sizes. Dilated convolution expands the receptive field by introducing a dilation rate, thereby increasing the receptive field of the network without increasing the number of parameters, which helps to extract multi-scale features.
[0083] To fuse these temporary feature maps, a content-aware upsampling method is used to expand the low-scale feature maps.
[0084] After multiple convolutions and upsamplings, the data with the largest size needs to pass through a CA-3D attention convolution module, whose main purpose is to enhance the feature expression ability of the model and convert the temporary entity channels to the standard number, that is, three channels. In the final output, N represents the size of the output volumetric texture, which is generally a multiple of 2 (for example, 128, 256, 512, etc.).
[0085] As a preferred embodiment, the first hybrid dilated convolution block and the second hybrid dilated convolution block have the same structure, both including three convolutional layers;
[0086] Among them, the convolutional kernels of the first and third convolutional layers are 1×1×1, and the second convolutional layer is a convolutional layer with a dilation rate of 2.
[0087] As a specific embodiment, the specific form of the hybrid dilated convolution block is asFigure 3 As shown, the hybrid dilated convolution block includes three layers of convolution structures. The convolution kernel of the first layer of convolution is 1×1×1, the second layer is a convolution layer with a dilation rate of 2, and the structure of the third layer of convolution is the same as that of the first layer of convolution. Such a design keeps the number of convolution layers unchanged while reducing the number of parameters of the convolution kernel. At the same time, different dilation rates are used for consecutive convolution layers to prevent the grid effect. In addition, this structure can ensure that when the size of the input tensor is equal, the output of the hybrid dilated convolution block is consistent with that of the original convolution block, which greatly facilitates the integration and use of the hybrid dilated convolution block in the generator network architecture.
[0088] It should be noted that the dilation rate refers to the interval between each value in the convolution kernel. When the dilation rate is d (d≥1, d∈N*), the interval between each value in the convolution kernel is d - 1.
[0089] In order to highlight key information and improve the effectiveness of feature representation, a CA-3D attention convolution module is introduced in the generator in this embodiment to obtain attention weights in the three directions of length, width, and height of the volumetric texture and encode precise position information. As a preferred embodiment, the CA-3D attention convolution module includes a pooling layer, a feature concatenation convolution layer, a normalization activation layer, a feature decomposition layer, and an output layer connected in sequence;
[0090] The pooling layer is used to perform global average pooling on the input original feature map in the three directions of length, width, and height to obtain feature maps in the three directions;
[0091] The feature concatenation convolution layer is used to concatenate the feature maps in the three directions of length, width, and height, and use a convolution module with a convolution kernel of 1×1×1 to reduce the channel dimension to obtain a low-dimensional feature map;
[0092] The normalization activation layer is used to perform a normalization operation on the low-dimensional feature map and a non-linear operation on the normalized feature map;
[0093] The feature decomposition layer is used to perform a segmentation operation along the spatial dimension to restore the feature maps in the three directions, use a convolution module of 1×1×1 to restore the channel dimension, and obtain the attention weights of the feature maps in the three directions through an activation function;
[0094] The output layer is used to weight the original feature map and the attention weights of the feature maps in the three directions, and convert the channel dimension into three dimensions through a convolution module of 1×1×1.
[0095] As a specific embodiment, the specific structure of the CA-3D attention convolution module is as Figure 4 shown. First, the input feature map is globally averaged in the three directions of length, width, and height to obtain feature maps in the three directions.
[0096] In the second step, the feature maps in three directions are concatenated together, and the concatenated feature maps are fed into a shared convolutional module with a convolution kernel of 1×1×1 to reduce its channel dimension to C / r, obtaining a feature map of size C / r×1×1×(D+H+W).
[0097] In the third step, a normalization operation is performed, and a non-linear operation is carried out on the normalized feature map through the Sigmoid activation function.
[0098] In the fourth step, a split operation is performed along the spatial dimension to restore the feature maps in three directions. Meanwhile, the channel dimension is restored using a 1×1×1 convolutional module, and the attention weights of the feature maps in three directions are obtained through the Sigmoid activation layer.
[0099] In the fifth step, the attention weights of the feature maps in three directions obtained are used to perform multiplicative weighting calculation on the original feature maps to obtain the feature maps with attention weights in the length, width, and height directions.
[0100] In the sixth step, the channel dimension of this feature map is converted to the standard three dimensions through a 1×1×1 convolutional module.
[0101] The above CA attention mechanism takes into account channel, spatial relationships, and long-range dependencies, and has a relatively low computational complexity. By introducing the CA module into the generator, the problem of discontinuous distortion in synthetic volume textures in traditional volume texture generation methods can be effectively solved.
[0102] As a preferred embodiment, as Figure 5 shown, the upsampling module includes an upsampling layer, a three-dimensional convolutional layer, a BN layer, and a Relu activation layer connected in sequence;
[0103] Among them, the upsampling module includes a prediction unit and a feature recombination unit;
[0104] The prediction module compresses the number of channels through 1×1 convolution, converts the number of channels using content encoding, and performs an expansion operation on the channel dimension in the spatial dimension to generate an upsampling kernel. The upsampling kernel is normalized to keep the original feature content unchanged, obtaining an output feature map;
[0105] The feature recombination module is used to map the local information of the output feature map to the input feature map, complete the recombination of the feature map, and output the upsampling output value.
[0106] As a specific embodiment, the upsampling module of the present invention is improved based on the CARAFE module and adopts the method of nearest neighbor upsampling.
[0107] The content-aware upsampling module provided in this embodiment is divided into a prediction unit for the upsampling kernel and a feature recombination unit.
[0108] Among them, the working process of the prediction unit of the upsampling kernel is as follows:
[0109] Assume that the upsampling ratio is σ, given an input feature map with a shape of H×W×C, and the size of the upsampling kernel is K up ×K up . In the prediction module of the upsampling kernel, for the input feature map, first use a 1×1 convolution to compress the number of channels to C m ,C m = C / r. In the rock texture generation task of this embodiment, the purpose of the generator is to generate volume data. Therefore, the number of channels of the input feature map is adjusted to C m = C / (2*2*2), and the result of the number of channels needs to be a positive integer to achieve the upsampling of small-scale data.
[0110] Second step, perform content encoding, convert the number of channels to σ 3 ×K up 3 , and then expand the channel dimension in the spatial dimension to obtain an upsampling kernel with a shape of σH×σW×σD×K up 3 .
[0111] Third step, perform a normalization operation to avoid changing the average value of the feature map and thus ensure the invariance of the original content.
[0112] In the feature recombination module, for each position in the output feature map, map it to the input feature map, extract a region of K up ×K up , and take the dot product with the predicted upsampling kernel at this point to obtain the output value of the upsampling.
[0113] Next, the structure of the discriminator of the rock texture synthesis model will be described in detail in combination with Figure 6 .
[0114] The discriminator of the GAN network is constructed based on multiple scales, and the structure is as Figure 6 shown. The discriminator structure of each scale can refer to the discriminator (STD) in STS-GAN. As a preferred embodiment, each scale of the discriminator is used to judge the authenticity of the image at the corresponding scale, including:
[0115] Each scale of the discriminator is used to perform texture discrimination on the slice data extracted after slicing the volume data at a preset scale and the real 2D sample data;
[0116] Among them, the slice data is obtained by sampling the volume texture generated by the generator, and volume textures of different scales are formed for slicing; the input real 2D samples are adjusted to different scales through sampling operations to achieve discrimination at different scales, and the real 2D sample data and the slice data in each discriminator have the same resolution.
[0117] Furthermore, the input 2D samples are adjusted to different scales by sampling operations on the input images, and then the corresponding discrimination operations are run at each scale. To increase the diversity of real samples, multiple texture blocks of predefined sizes are randomly cropped from the real samples, and then the sizes of these blocks are adjusted to the same resolution to provide multi-scale "real" textures for the discriminator.
[0118] To better train the model, as a preferred embodiment, the rock image data is collected and preprocessed to obtain a sample data set, including:
[0119] Obtain real rock images with a preset resolution and screen the rock images;
[0120] Manually crop the texture regions in the screened rock images;
[0121] Resample the cropped images to generate texture pictures of different resolutions, and generate a sample data set of rock texture images.
[0122] As a specific embodiment, part of the real rock image data collected in the present invention is from the rock texture pictures in some publicly available texture data sets on the Internet and some texture pictures in papers and books; the other part is from the high-definition rock images publicly available in the National Demonstration Center for Experimental Geology Teaching of Chengdu University of Technology and the Yifu Museum of China University of Geosciences.
[0123] In some embodiments, as Figure 7 shown, the preprocessing of the real rock image data includes:
[0124] The first step is to select suitable real rock images for cropping from the real rock images with a pixel resolution of more than 1000 pixels collected manually, that is, Figure 7 the leftmost image in
[0125] The second step is to manually crop the large texture regions in the image. The cropped image has a resolution of about 600×600 pixels, that is, Figure 7 the middle image in
[0126] The third step is to resample the image to form texture pictures with resolutions of 128 pixels, 256 pixels, and 512 pixels, as Figure 7For the rightmost image, store the processed picture in a folder for subsequent model training.
[0127] As a preferred embodiment, use the rock texture image sample dataset to train the rock body texture synthesis model to obtain a trained rock body texture synthesis model, including:
[0128] Set the learning rates of the generator and discriminator respectively, and use the AdamW optimizer for training;
[0129] Introduce a learning rate decay strategy during the training process. After every preset number of training times, halve the learning rate to optimize the training effect of the model.
[0130] In some embodiments, for the multi-scale features of the rock texture image, set the learning scale of the generator to 5, the generator learning rate to 5e -4 , the discriminator learning rate to 3e -4 , use the AdamW optimizer during the training process. It should be noted that the AdamW optimizer is a variant of the Adam optimizer, which solves the problem of insufficient weight decay that may occur in the Adam optimizer during the training process by adding weight decay. When calculating the gradient update, the AdamW optimizer directly adds the weight decay term to the gradient instead of performing the weight decay operation after updating the parameters. This can more accurately control the degree of weight decay and avoid the situation of over-decay or under-decay. The slice resolution of the default output samples during training is the same as the resolution of the input samples. At the same time, introduce a learning rate decay strategy, set the preset number of training times (such as 6000 times), and halve the learning rate every time 6000 training times are reached to optimize the training effect of the model and obtain the optimal parameter model. The specific training parameters are shown in Table 1.
[0131] Table 1
[0132]
[0133] As a specific embodiment, the visualization display of the generated rock texture synthesis map includes: input the rock texture image into the trained model, generate numpy data format and use OpenGL for visualization display, and finally obtain the rock body texture synthesis effect.
[0134] To verify the actual effect of the present invention, as a specific embodiment, select various rock texture images from the constructed rock texture dataset and input them into the trained model for experiments. Figure 8 The rock texture images used in the experiment are shown. There are a total of five rock images in the figure.
[0135] Perform experiments on the above 5 kinds of data. The schematic diagram of the volume texture synthesis result is asFigure 9 As shown in the figure, the first column from top to bottom are the real texture images of Sanbao granite, plagioclase, augite, green peridotite, and shell limestone. The second column is the result of volume texture synthesis, and the third to sixth columns are the sliced effects in three orthogonal directions and along the 45° diagonal direction of the solid texture. It can be seen from the figure that the method proposed in the present invention can better learn the texture feature distribution on the two-dimensional texture image and extend it to the three-dimensional space, showing good consistency in different directions of the three-dimensional space.
[0136] As Figure 10 shown, this embodiment also provides a rock mass texture synthesis system 1000 based on CA-GAN, including:
[0137] A data acquisition module 1001, configured to collect rock image data, preprocess the rock image data, and obtain a rock texture image sample data set;
[0138] A model construction module 1002, configured to introduce a CA attention mechanism and a content-aware upsampling module into the architecture of the generative adversarial network, and construct a rock mass texture synthesis model based on CA-GAN;
[0139] A model training module 1003, configured to use the rock texture image sample data set to train the rock mass texture synthesis model, and obtain a well-trained rock mass texture synthesis model;
[0140] An image synthesis module 1004, configured to generate a rock mass texture synthesis map through the well-trained rock mass texture synthesis model, and perform visual display on the rock texture synthesis map.
[0141] The method and system for rock mass texture synthesis based on CA-GAN proposed by the present invention combines the CA attention mechanism, the content-aware upsampling module with GAN. This method can effectively solve problems such as detail loss, texture repetition, insufficient local consistency, and poor global consistency in texture generation in the prior art. It can generate higher-quality, more realistic and diverse rock textures, adapt to a wider range of application scenarios, and has significant advantages especially when dealing with complex textures and large-scale scenarios.
[0142] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.
Claims
1. A rock texture synthesis method based on CA-GAN, characterized in that: include: Collect rock image data, pre-process the rock image data, and obtain a rock texture image sample data set; The CA attention mechanism and content-aware upsampling module are introduced into the architecture of the generative adversarial network to build a rock texture synthesis model based on CA-GAN. The rock body texture synthesis model is trained using a rock texture image sample data set to obtain a fully trained rock body texture synthesis model; A rock texture synthesis image is generated by training a complete rock texture synthesis model, and the rock texture synthesis image is visualized.
2. The rock texture synthesis method based on CA-GAN according to claim 1, characterized in that: The CA attention mechanism and the content-aware upsampling module are introduced into the architecture of the generative adversarial network to construct an initial rock texture synthesis model based on the generative adversarial network, including: A hybrid atrous convolution layer, CA attention mechanism and content-aware upsampling module are introduced into the generator of the generative adversarial network; feature information of different sizes in rock image data is extracted through the hybrid atrous convolution layer; the size of the input feature map is enlarged through the content-aware upsampling module; the CA attention mechanism is used to obtain attention weights in the length, width and height directions of the rock texture, encode the texture position information, and restore the detail information of the image to generate a rock texture with rich details; In the discriminator of the generative adversarial network, multiple discriminators of different scales are used. The discriminator of each scale is used to judge the authenticity of the image at the corresponding scale. The judgment results of different scales are integrated to obtain a comprehensive evaluation of the generated image.
3. The rock texture synthesis method based on CA-GAN according to claim 2 is characterized in that: The structure of the generator of the rock texture synthesis model includes: a multi-scale input module, a first mixed hole convolution block, a second mixed hole convolution block, an upsampling module and a CA-3D attention convolution module; The input module of each scale is connected to the corresponding first mixed void convolution block to input multi-scale random volume noise so that the model outputs different textures; The first mixed atrous convolution block is used to process the input noise and generate temporary feature maps of different scales; The output end of the first hybrid atrous convolution block of the smallest size is connected to the input end of the upsampling module, the output ends of the first hybrid atrous convolution blocks of the remaining sizes are connected to the second hybrid atrous convolution block, and the output end of the second hybrid atrous convolution block is connected to the input end of the upsampling module; The upsampling module expands the temporary feature map through a content-aware upsampling method; Except for the temporary feature map of the smallest size, the temporary feature maps of the remaining sizes and the output feature maps of the upsampling modules of the adjacent secondary sizes are used as input data and input into the second hybrid dilated convolution block, which is used to fuse the temporary feature maps of different sizes; The output of the second hybrid atrous convolutional block of the largest size is connected to the input of the CA-3D attention convolutional module; The CA-3D attention convolution module is used to obtain the attention weights of the rock texture in the length, width, and height directions, convert the entity channel to three channels, and encode the position information.
4. The rock texture synthesis method based on CA-GAN according to claim 3 is characterized in that: The first mixed dilated convolution block and the second mixed dilated convolution block have the same structure, and both include three convolution layers; Among them, the convolution kernels of the first and third convolutional layers are 1×1×1, and the second convolutional layer is a convolutional layer with a dilation rate of 2.
5. The rock texture synthesis method based on CA-GAN according to claim 3 is characterized in that: The CA-3D attention convolution module includes a pooling layer, a feature convolution layer, a normalized activation layer, a feature decomposition layer and an output layer connected in sequence; The pooling layer is used to perform global average pooling on the input original feature map in the length, width, and height directions to obtain feature maps in three directions; The feature concatenation convolution layer is used to concatenate feature maps in length, width, and height, and use a convolution module with a convolution kernel of 1×1×1 to reduce the channel dimension to obtain a low-dimensional feature map. The normalized activation layer is used to normalize the low-dimensional feature map and perform nonlinear operations on the normalized feature map; The feature decomposition layer is used to perform segmentation operations along the spatial dimension to restore the feature maps in three directions, use a 1×1×1 convolution module to restore the channel dimension, and obtain the attention weights of the feature maps in three directions through the activation function; The output layer is used to weight the original feature map and the attention weights of the feature maps in three directions, and the channel dimension is converted into three dimensions through a 1×1×1 convolution module.
6. The CA-GAN-based rock texture synthesis method according to claim 3, characterized in that: The upsampling module includes an upsampling layer, a three-dimensional convolution layer, a BN layer and a Relu activation layer connected in sequence; Among them, the upsampling layer includes a prediction unit and a feature recombination unit; The prediction module compresses the number of channels through 1×1 convolution, converts the number of channels using content coding, and expands the channel dimension in the spatial dimension to generate an upsampling kernel, which is normalized to keep the original feature content unchanged to obtain the output feature map; The feature recombination module is used to map the local information of the output feature map to the input feature map, complete the recombination of the feature map, and output the upsampled output value.
7. The rock texture synthesis method based on CA-GAN according to claim 2, characterized in that: The discriminator of each scale is used to judge the authenticity of the image at the corresponding scale, including: The discriminator at each scale is used to perform texture identification between the slice data extracted after slicing the volume data at a preset scale and the real 2D sample data; The slice data is obtained by sampling the volume texture generated by the generator to form volume textures of different scales for slicing; the input real 2D samples are adjusted to different scales through sampling operations to achieve identification at different scales, and the real 2D sample data and the slice data in each discriminator have the same resolution.
8. The rock texture synthesis method based on CA-GAN according to claim 1, characterized in that: The rock image data is collected and preprocessed to obtain a sample data set, including: Acquire real rock images with preset resolution and screen the rock images; Manually cropping the texture area in the screened rock image; The cropped image is resampled to generate texture images of different resolutions and a rock texture image sample dataset.
9. The rock texture synthesis method based on CA-GAN according to claim 1, characterized in that: The rock body texture synthesis model is trained using a rock texture image sample data set to obtain a fully trained rock body texture synthesis model, including: Set the learning rates of the generator and discriminator respectively, and use the AdamW optimizer for training; The learning rate decay strategy is introduced into the training process. After each preset number of training times, the learning rate is halved to optimize the training effect of the model.
10. A rock texture synthesis device based on CA-GAN, characterized in that: include: A data acquisition module is used to collect rock image data, pre-process the rock image data, and obtain a rock texture image sample data set; The model building module is used to introduce the CA attention mechanism and content-aware upsampling module into the architecture of the generative adversarial network to build a rock texture synthesis model based on CA-GAN; A model training module is used to train the rock body texture synthesis model using a rock texture image sample data set to obtain a fully trained rock body texture synthesis model; The image synthesis module is used to generate a rock texture synthesis image by training a complete rock texture synthesis model, and to visualize the rock texture synthesis image.