A weight training method and system for generating geological lithological texture models based on LoCon
By incorporating LoCon's low-rank matrix into the LDM and U-Net models, and combining it with the professional knowledge of geologists, high-quality geological lithology texture maps can be generated quickly on ordinary graphics cards. This solves the problems of low texture generation efficiency and high threshold in the geological industry, and meets professional needs.
Patent Information
- Application Number
- CN202511096069.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2045-08-06
AI Technical Summary
The geological industry lacks efficient texture generation methods. Existing technologies are unable to quickly generate material textures that conform to geological lithological characteristics. Furthermore, the high barrier to entry for geologists in the digitization process results in insufficient material libraries and rough digital display effects that fail to meet professional needs.
We employ the Latent Diffusion Model (LDM) as the weight training framework for the geological lithology texture model. Combining the denoising core of the U-Net model and the low-rank matrix of the LoCon model, we extend the convolutional layers and incorporate residual blocks. We then use professional knowledge language to describe and generate textures that conform to geological characteristics, thereby reducing the operational threshold.
It can quickly generate high-quality geological lithology texture maps on a small number of samples and ordinary graphics card computers, reducing the technical threshold for geologists, improving texture generation efficiency and quality, and meeting professional needs.
Smart Images

Figure CN120852623B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of geological lithological material texture generation technology, specifically to a weight training method and system for generating geological lithological material texture models based on LoCon. Background Technology
[0002] Currently, the geological industry faces significant bottlenecks in its digital transformation, particularly in the representation of geological lithology. Geologists struggle to quickly obtain material data that meets their professional needs, resulting in crude digital displays that fail to satisfy reporting, teaching, or engineering applications. There is currently no dedicated material library for geological lithology characteristics, either domestically or internationally. The textures used to generate stratigraphic materials are generally based on surface sand, gravel, and soil textures, while complete rock strata representations are extremely scarce. Existing geological materials are typically obtained through image acquisition from rock sample photographs or rock scanning, requiring substantial manual labor to achieve clean textures. Furthermore, the limited availability of geological lithology samples, coupled with scarce and undiversified data, hinders the establishment of a geological material library.
[0003] Generating the textures required for geology necessitates geologists with expertise in lithological characteristics. However, in the current digital field, texture generation typically employs manual drawing or automated rule-based drawing, which are too technically demanding for geologists, complex to operate, and produce results that deviate from the intended objectives. In AI applications, while image generation technology is relatively mature, some studies have attempted to fine-tune the generation of natural textures from publicly available image datasets. However, these studies do not target geological lithological characteristics, and the generated textures are still based on natural surface textures.
[0004] Currently, the core problem facing the geological industry in the direction of digitalization is the lack of an efficient method for generating textures of lithological materials. This situation stems not only from the limitations of the technology itself, but also from the fact that geologists cannot directly participate in the digitalization process, nor can they translate their professional knowledge into digital results. Summary of the Invention
[0005] To address the problems existing in current technologies, the purpose of this invention is to propose a weight training method and system for generating geological lithological texture models based on LoCon. The aim is to enable the generation of geological lithological texture models using a small number of samples and ordinary consumer-grade graphics cards on a computer, based on the weights trained according to this invention. Professionals can then use these weights, described using specialized technical language, to efficiently generate textures that conform to geological characteristics, filling gaps in the geological industry's material library and lowering the technical threshold for geologists to participate in digitalization. This invention is based on a pre-trained geological lithological texture model. It extends the convolutional layer in the LoCon model and integrates this convolutional layer into the convolutional weights of the residual blocks in the U-Net model, enabling it to quickly adapt to the specific characteristics of geological lithological textures and generate high-quality two-dimensional texture maps.
[0006] This invention provides a weight training method for generating geological lithological material texture models based on LoCon, comprising the following steps:
[0007] Step S100: The Latent Diffusion Model (LDM) is used as the large model framework for weight training and texture generation of the geological lithology texture model. The pre-trained weight matrix is loaded using the denoising core of the U-Net model. The convolutional layer is extended in the LoCon model and integrated into the convolutional weights of the residual block of the U-Net model. The low-rank matrix of the LoCon model is added to the pre-trained weight matrix in the U-Net model to build the geological lithology texture model.
[0008] Step S200: Obtain sample images of geological lithological material textures, perform texture processing, feature annotation, and data augmentation on the sample images to form a training set;
[0009] Step S300: Input the training set and train the generated geological lithology texture model according to the pre-trained weight matrix of the low-rank matrix with LoCon model added, until the preset training requirements are met.
[0010] Step S400: Output the trained weights, load the diffusion model LDM large model framework to generate geological lithological material textures, adjust the training parameters according to the generated texture results, and re-execute step S300 until the preset texture result requirements are met. Use the diffusion model LDM large model framework and the trained weight matrix to create the interface API for generating geological lithological material textures.
[0011] Step S300 specifically includes the following steps:
[0012] Step S301, Training Preparation: Load the LDM large model framework and configure the training parameters;
[0013] Step S302, Start Training: Check the parameter update status of the LoCon model by the average loss rate of each training round. Adjust the training parameters according to the convergence status reflected by the parameters of the LoCon model, and retrain until the convergence speed meets the requirements.
[0014] Preferably, in step S100, the pre-trained model uses a variational autoencoder (VAE), uses Clip as the text decoder, and uses cosion with restart as the scheduler.
[0015] Preferably, in step S100, after adding the low-rank matrix of LoCon to the pre-trained weight matrix of U-Net in the pre-trained model, the original weights are frozen.
[0016] Preferably, in step S100, all components are loaded using the pipe encapsulated in the pipe.
[0017] Preferably, in step S200, images are collected from the network under the guidance of geological personnel, or sample photos are provided, and after texturization processing, the images are made to ensure that they reflect the target features.
[0018] Preferably, in step S200, the image is processed using an image editing tool to perform pure texturing. The processed image does not contain spatial orientation, perspective, light and shadow or object outline of three-dimensional information, but only reflects the color and features of the material surface.
[0019] Preferably, in step S200, the samples are labeled with features by text description or keywords, and a labeling tool plugin is used to generate a labeling file and corresponding icon name labels to generate a label set.
[0020] Preferably, the feature annotation includes trigger words and additional descriptive words, with the trigger words being placed at the beginning of the label set as the main identifier for training.
[0021] Preferably, the initial tag set is cleaned and corrected by deleting feature tags related to the target ontology in the training set; and the order of the remaining tags, except for the trigger words, is randomly shuffled.
[0022] Preferably, the training set is reviewed and adjusted in batches, and erroneous labels and labels irrelevant to the target in the training set are deleted in batches; for images with common characteristics, descriptive labels of common characteristics are added in batches.
[0023] Preferably, for individual image elements not recognized by the labeling tool, corresponding labels are added to the individual images through manual comparative analysis.
[0024] Preferably, in step S200, data augmentation specifically includes: repeatedly copying a single feature sample in the training set, feature classification, multi-dimensional changes to the sample image, and shuffling a random seed.
[0025] Preferably, feature classification includes: performing multi-bucket classification on images in the training set, cropping the texture of images with a ratio within a preset range, and grouping them into the same bucket, with each bucket representing a set of images of a certain resolution.
[0026] Preferably, the multidimensional changes of the sample include: performing color space adjustment and geometric transformation operations on the sample image. The color space adjustment includes brightness transformation, contrast adjustment and hue transformation. The geometric transformation operations include horizontal flipping and vertical flipping to simulate the color characteristics and symmetry variants of the sample under different deposition environments.
[0027] Preferably, in step S301, the training parameters set include the total number of steps, U-Net learning rate, Text Encoder learning rate, network dimension, and optimizer. After freezing the original pre-trained weights, the number of training epochs is set.
[0028] Preferably, in step S302, the convergence status of the LoCon model is observed by checking the average loss rate of each training epoch. If the convergence is too fast, the learning rate is reduced; if the convergence is too slow, the step size parameter is increased. The convergence speed is judged by comparing the number of training epochs at which convergence is achieved with a preset training epoch threshold. If the number of training epochs at which convergence is achieved is higher than the upper limit of the preset training epoch threshold range, the convergence speed is considered too slow; if the number of training epochs at which convergence is achieved is lower than the lower limit of the preset training epoch threshold range, the convergence speed is considered too fast.
[0029] Preferably, in step S400, the trained weights are output, the diffusion model LDM large model framework is loaded to generate geological lithological material textures, and the training parameters are adjusted based on the generated texture results, specifically including:
[0030] If the number of times a geological lithological texture model is generated exceeds a preset number, the similarity of the generated geological lithological texture effect is lower than the first preset threshold, or an irrelevant background is generated, then overfitting occurs, and the learning rate is reduced.
[0031] If the geological lithological texture model generates a geological lithological texture effect that does not contain some features or has a similarity to some features that is lower than the second preset threshold, then the number of repetitions of a single feature sample is increased.
[0032] If the generated geological lithology texture model produces a texture effect that has no texture or has a similarity of 0 with the expected texture, then underfitting occurs. In this case, increase the total number of steps and the learning rate, increase the Rank parameter in the LoCon model, decrease the Alpha parameter in the LoCon model, and increase the number of repetitions for a single feature sample.
[0033] Preferably, the reasoning in step S400:
[0034] The present invention also provides a weight training system for generating geological lithological texture models based on LoCon, including a processor, the processor being able to execute a computer program, the computer program being able to implement the above-described weight training method for generating geological lithological texture models based on LoCon.
[0035] This invention proposes a weight training method for a geological lithological material texture model based on LoCon, which has significant advantages in technical performance compared to conventional techniques (such as manual drawing, automated rule drawing, or AI generation methods based on natural surface textures). Compared to existing technologies, this invention has the following beneficial effects:
[0036] (1) This invention achieves efficient generation of high-quality geological lithological textures through lightweight fine-tuning and highly adaptable techniques. Compared to the inefficiency of conventional techniques that rely on a large number of samples or complex computing resources to generate textures, this invention can quickly generate high-quality two-dimensional texture maps that conform to geological lithological characteristics using a small number of samples (20-50 images) on a standard consumer-grade graphics card computer. This effect stems from the use of LoCon (LoRA for Convolution Network) technology, which adds a low-rank update matrix to the pre-trained Latent Diffusion Model (LDM) framework, allowing for fine-tuning of only a few parameters to adapt to the specific features of geological lithological textures. By extending the convolutional layer in the LoCon model and integrating it into the convolutional weights of the residual blocks in the U-Net model, this LoCon extended convolutional layer design further enhances the ability to capture low-level detail features, ensuring that the generated texture is closer to the real geological requirements in terms of detail representation.
[0037] (2) This invention achieves the technical effect of allowing geologists to directly participate in digitization through professional knowledge language descriptions and low-threshold operational techniques. In conventional techniques, texture generation usually requires geologists to master complex digitization tools (such as manual drawing or automated rule drawing), resulting in high operational barriers and results that deviate from the target. This invention, by combining LoCon technology and CLIP text decoder, allows geologists to input professional knowledge with simple text descriptions or keywords, driving the model to generate textures that conform to lithological characteristics. The combination of LoCon's low-rank matrix and CLIP's text encoding capabilities enables the model to efficiently parse geological feature descriptions and generate textures that meet the requirements. This technical approach allows geologists to participate in the digitization process without needing to learn complex tools in depth, significantly reducing the technical threshold.
[0038] (3) This invention improves the model generation effect and training efficiency through a comprehensive training method. This invention enhances the model's ability to capture low-level features by extending the convolutional layer in the LoCon model framework and integrating it into the residual blocks of U-Net, thereby capturing the features of geological lithology texture in a better way. In the data processing stage, the model avoids bias and enhances the model's understanding of geological features by using a more efficient combination of automatic and manual labeling. Attached Figure Description
[0039] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.
[0040] Figure 1 This is a flowchart of a weight training method for generating a geological lithology texture model based on LoCon, according to an embodiment of the present invention.
[0041] Figure 2 This is a flowchart illustrating the creation process of generating a geological lithological material texture model according to an embodiment of the present invention.
[0042] Figure 3 This is an embodiment of the U-Net internal framework of the present invention;
[0043] Figure 4 The original image of a training sample image according to one embodiment of the present invention;
[0044] Figure 5 for Figure 4 The image obtained by randomly horizontally flipping the original image;
[0045] Figure 6 for Figure 4 The image obtained by randomly rotating the original image;
[0046] Figure 7 for Figure 4 The image obtained after color dithering of the original image;
[0047] Figure 8 for Figure 4 The image obtained after randomly converting the original image to grayscale;
[0048] Figure 9 for Figure 4 The image obtained by randomly cropping the original image;
[0049] Figure 10 This is a training log for generating a geological lithology texture model according to an embodiment of the present invention;
[0050] Figure 11 This is a demonstration of the accuracy of the training rounds in one embodiment of the present invention. Detailed Implementation
[0051] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0052] like Figure 1 As shown, this invention provides a weight training method for generating geological lithological material texture models based on LoCon, including the following steps:
[0053] Step S100: The Latent Diffusion Model (LDM) is used as the large model framework for weight training and texture generation of the geological lithology texture model. The pre-trained weight matrix is loaded using the denoising core of the U-Net model. The convolutional layer is extended in the LoCon model and integrated into the convolutional weights of the residual block of the U-Net model. The low-rank matrix of the LoCon model is added to the pre-trained weight matrix in the U-Net model to build the geological lithology texture model.
[0054] A core component of the LDM large model framework is called the U-Net model, which is part of the framework. The LoCon model is a technical means. The LoCon model is attached to the U-Net model through matrix addition and channel splicing, so that the U-Net model can play the additional functions provided by the LoCon technical means.
[0055] Load the LoCon class onto the weight matrix of the U-Net model, perform weight matrix addition to obtain the weight matrix of the U-Net model with the LoCon model attached, execute the training program, and obtain the trained LoCon model.
[0056] Step S200: Obtain sample images of geological lithological material textures, perform texture processing, feature annotation, and data augmentation on the sample images to form a training set;
[0057] Step S300: Input the training set and train the generated geological lithology texture model according to the pre-trained weight matrix of the low-rank matrix with LoCon model added, until the preset training requirements are met.
[0058] Step S400: Output the trained weights, load the diffusion model LDM large model framework to generate geological lithological material textures, adjust the training parameters according to the generated texture results, and re-execute step S300 until the preset texture result requirements are met. Use the diffusion model LDM large model framework and the trained weight matrix to create the interface API for generating geological lithological material textures.
[0059] Step S300 specifically includes the following steps:
[0060] Step S301, Training Preparation: Load the LDM large model framework and configure the training parameters;
[0061] Step S302, Start Training: Check the parameter update status of the LoCon model by the average loss rate of each training round. Adjust the training parameters according to the convergence status reflected by the parameters of the LoCon model, and retrain until the convergence speed meets the requirements.
[0062] According to a specific embodiment of the present invention, in step S100, the pre-trained model adopts a variational autoencoder (VAE), uses Clip as a text decoder, and uses cosion with restart as a scheduler.
[0063] According to a specific embodiment of the present invention, in step S100, after adding the low-rank matrix of LoCon to the pre-trained weight matrix of U-Net in the pre-trained model, the original weights are frozen.
[0064] According to a specific embodiment of the present invention, in step S100, all components are loaded using the pipe-encapsulated pipeline.
[0065] According to a specific embodiment of the present invention, step S100 specifically integrates the convolutional layer into the convolutional weights of the residual block of the U-Net model in the following manner:
[0066] Parameters are extracted from the convolutional layers extended from the LoCon model to obtain the convolutional kernel parameter matrix WLoCon, and the weight matrix WU−Net of the target convolutional layer in the residual block of the U-Net model is obtained at the same time.
[0067] Add WLoCon to WU−Net element by element, that is, the new convolutional weight matrix Wnew=WU−Net+WLoCon, to achieve the integration of the convolutional layer;
[0068] Next, the low-rank matrix MLoCon of the LoCon model is added to the pre-trained weight matrix Wpretrained in the U-Net model by matrix addition, i.e., Wfinal=Wpretrained+MLoCon, to obtain the initial weight matrix, and then the geological lithology material texture model is built.
[0069] According to a specific embodiment of the present invention, in step S200, images are collected from the network under the guidance of geological personnel, or sample photos are provided, and after texturing processing, the images are made to ensure that they reflect the target features.
[0070] According to a specific embodiment of the present invention, in step S200, an image is processed by using an image editing tool to perform pure texturing. The processed image does not contain spatial orientation, perspective, light and shadow or object outline of three-dimensional information, but only reflects the color and features of the material surface.
[0071] According to a specific embodiment of the present invention, in step S200, the sample is characterized by text description or keywords, and a labeling tool plugin is used to generate a labeling file and corresponding icon name labeling to generate a label set.
[0072] According to a specific embodiment of the present invention, the feature annotation includes trigger words and additional descriptive words, wherein the trigger words are placed at the beginning of the label set as the main identifier for training.
[0073] According to a specific embodiment of the present invention, the initial tag set is cleaned and corrected by deleting feature tags related to the target ontology in the training set; and the order of the remaining tags, except for the trigger words, is randomly shuffled.
[0074] For example, to generate granite texture, the feature labels for granite are "quartz grains" and "cross-bedding." When organizing the labels, these two granite feature labels should be deleted. This is because if the training is performed with these labels, the default prompt for granite will be that these two features must be included in order to generate granite texture. This would solidify the "granite" label in the training process, causing the system to automatically learn the features of the granite in the current sample.
[0075] According to a specific embodiment of the present invention, the training set is reviewed and adjusted in batches, and erroneous labels and labels irrelevant to the target in the training set are deleted in batches; for images with common characteristics, descriptive labels of common characteristics are added in batches.
[0076] Common features, such as certain granite photos in the training set, were all taken in low light or at night, resulting in complex lighting and shadows, and even obvious reflections of mobile phones on the rocks. These are not what is needed for training. If the common and unique features of these photos, such as "night" and "dappled light and shadow," are not explicitly included in the label, and the shadow is not clearly defined as unique to these photos, the training will default to the shadow feature as a granite feature. This may lead to the generation of textures containing mobile phone reflections when generating granite later. Therefore, incorrect common features should be removed.
[0077] According to a specific embodiment of the present invention, for single image elements not recognized by the marking tool, corresponding labels are added to the single image through manual comparison and analysis.
[0078] According to a specific embodiment of the present invention, in step S200, data augmentation specifically includes: repeatedly copying a single feature sample in the training set, feature classification, multi-dimensional changes to the sample image, and shuffling a random seed.
[0079] The individual feature samples in the training set are repeatedly copied. Originally there were 5 images, but this was repeated 4 times, resulting in 20 images. These 20 images are then processed. The reason for copying is that some feature maps have very little data. For example, there is only one feature map of the bedding structure of granite. By copying them, different bedding structure feature map variants are generated through color space and geometric transformations, avoiding overfitting during training due to a single, overly simplistic image.
[0080] According to a specific embodiment of the present invention, feature classification includes: performing multi-bucket classification on images in the training set, cropping the texture of images with a ratio within a preset range, and grouping them into the same bucket, with each bucket representing a set of images of a certain resolution.
[0081] According to a specific embodiment of the present invention, the multidimensional changes of the sample include: performing color space adjustment operations and geometric transformation operations on the sample image, wherein the color space adjustment includes brightness transformation, contrast adjustment and hue transformation; and the geometric transformation operations include horizontal flipping and vertical flipping to simulate the color characteristics and symmetry variants of the sample under different deposition environments.
[0082] According to a specific embodiment of the present invention, in step S301, the training parameters set include the total number of steps, U-Net learning rate, Text Encoder learning rate, network dimension, and optimizer. After freezing the original pre-trained weights, the number of training epochs is set.
[0083] According to a specific embodiment of the present invention, in step S302, the convergence status of the LoCon model's parameter update state is viewed through the average loss rate of each training epoch. If the convergence is too fast, the learning rate is reduced; if the convergence is too slow, the step size parameter is increased. The convergence speed is judged by comparing the training epoch at which convergence is achieved with a preset training epoch threshold. If the training epoch at which convergence is achieved is higher than the upper limit of the preset training epoch threshold range, the convergence speed is considered too slow; if the training epoch at which convergence is achieved is lower than the lower limit of the preset training epoch threshold range, the convergence speed is considered too fast.
[0084] According to a specific embodiment of the present invention, in step S400, the trained weights are output, the diffusion model LDM large model framework is loaded to generate geological lithological material textures, and the training parameters are adjusted according to the generated texture results, specifically including:
[0085] If the number of times a geological lithological texture model is generated exceeds a preset number, the similarity of the generated geological lithological texture effect is lower than the first preset threshold, or an irrelevant background is generated, then overfitting occurs, and the learning rate is reduced.
[0086] If the geological lithological texture model generates a geological lithological texture effect that does not contain some features or has a similarity to some features that is lower than the second preset threshold, then the number of repetitions of a single feature sample is increased.
[0087] If the generated geological lithology texture model produces a texture effect that has no texture or has a similarity of 0 with the expected texture, then underfitting occurs. In this case, increase the total number of steps and the learning rate, increase the Rank parameter in the LoCon model, decrease the Alpha parameter in the LoCon model, and increase the number of repetitions for a single feature sample.
[0088] Rank and Alpha are both parameters of the low-rank matrix in the LoCon model. Rank, or rank, represents the matrix dimension and determines the number of parameters in the final output weights and the model's expressive power. A small Rank means fewer parameters, focusing only on the most obvious original features, sacrificing some features; a large Rank means more parameters and richer features, but it may also learn unnecessary features such as shadows, while increasing performance overhead.
[0089] Alpha, or scaling factor, controls the scaling ratio of the original weights, thus affecting their functionality. The original weights, which form the base model of LDM, acquire the functionality of LoCon after being attached. A lower alpha value means fewer features are learned; a higher alpha value means more features are learned. The learning and training process involves learning new knowledge. Whether too few or too many features are learned, a balance is achieved by adjusting the parameters.
[0090] According to a specific embodiment of the present invention, step S100 specifically includes the following steps:
[0091] 1. Model initialization: Load the initial weight matrix `W_final` of the U-Net model, which contains the low-rank matrix of the LoCon model, into the residual block convolutional layer of the U-Net model as the initial values of the model parameters; freeze the parameters of the `W_pretrained` part (i.e. the original weights of U-Net), and only allow the LoCon parameters corresponding to `M_LoCon` (such as the `down` and `up` layers of the low-rank matrix) to participate in gradient updates.
[0092] 2. Forward propagation computation: Input training set sample images are converted into latent vectors through the encoding process of the LDM framework (VAE encoder); the latent vectors enter the U-Net model, and denoising is performed based on the initial weight matrix to generate predicted latent vectors; the predicted vectors are restored to the generated images by the VAE decoder, and the loss (such as L1 loss, perceptual loss) is calculated with the sample images.
[0093] 3. Backpropagation optimization: During backpropagation of the loss function, only the gradient of the LoCon parameter (`M_LoCon`) is calculated, and the gradient of the original weights of U-Net is set to zero. The optimizer (such as AdamW) updates the low-rank matrix parameters of LoCon according to the gradient, so that `M_LoCon` gradually fits the geological lithology characteristics, while `W_pretrained` remains unchanged.
[0094] 4. Iterative update logic: In each round of training, the `W_pretrained` part of the initial weight matrix is always used as the baseline, and the `M_LoCon` part is continuously adjusted through training; when the training reaches the preset number of rounds or the loss converges, the final `M_LoCon` parameters are combined with the fixed `W_pretrained` to form the complete weights after training.
[0095] The present invention also provides a weight training system for generating geological lithological texture models based on LoCon, including a processor, the processor being able to execute a computer program, the computer program being able to implement the above-described weight training method for generating geological lithological texture models based on LoCon.
[0096] Example 1
[0097] This invention provides a weight training method for generating geological lithological material texture models based on LoCon, comprising the following steps:
[0098] Step S100: The Latent Diffusion Model (LDM) is used as the large model framework for weight training and texture generation of the geological lithology texture model. The pre-trained weight matrix is loaded using the denoising core of the U-Net model. The convolutional layer is extended in the LoCon model and integrated into the convolutional weights of the residual block of the U-Net model. The low-rank matrix of the LoCon model is added to the pre-trained weight matrix in the U-Net model to build the geological lithology texture model.
[0099] Step S200: Obtain sample images of geological lithological material textures, perform texture processing, feature annotation, and data augmentation on the sample images to form a training set;
[0100] Step S300: Input the training set and train the generated geological lithology texture model according to the pre-trained weight matrix of the low-rank matrix with LoCon model added, until the preset training requirements are met.
[0101] Step S400: Output the trained weights, load the diffusion model LDM large model framework to generate geological lithological material textures, adjust the training parameters according to the generated texture results, and re-execute step S300 until the preset texture result requirements are met. Use the diffusion model LDM large model framework and the trained weight matrix to create the interface API for generating geological lithological material textures.
[0102] Step S300 specifically includes the following steps:
[0103] Step S301, Training Preparation: Load the LDM large model framework and configure the training parameters;
[0104] Step S302, Start Training: Check the parameter update status of the LoCon model by the average loss rate of each training round. Adjust the training parameters according to the convergence status reflected by the parameters of the LoCon model, and retrain until the convergence speed meets the requirements.
[0105] Example 2
[0106] This invention provides a weight training method for generating geological lithological material texture models based on LoCon, comprising the following steps:
[0107] Step S100: The Latent Diffusion Model (LDM) is used as the large model framework for weight training and texture generation of the geological lithology texture model. The pre-trained weight matrix is loaded using the denoising core of the U-Net model. The convolutional layer is extended in the LoCon model and integrated into the convolutional weights of the residual block of the U-Net model. The low-rank matrix of the LoCon model is added to the pre-trained weight matrix in the U-Net model to build the geological lithology texture model.
[0108] A core component in the LDM large model framework is called the UNet-model, which is part of the framework. The LoCon model is a technical means. The LoCon model is attached to the UNet model through matrix addition and channel concatenation, so that the UNet-model can play the additional functions provided by the LoCon technical means.
[0109] Load the LoCon class onto the weight matrix of the U-Net model, perform weight matrix addition to obtain the weight matrix of the U-Net model with the LoCon model attached, execute the training program, and obtain the trained LoCon model.
[0110] Step S200: Obtain sample images of geological lithological material textures, perform texture processing, feature annotation, and data augmentation on the sample images to form a training set;
[0111] Step S300: Input the training set and train the generated geological lithology texture model according to the pre-trained weight matrix of the low-rank matrix with LoCon model added, until the preset training requirements are met.
[0112] Step S400: Output the trained weights, load the diffusion model LDM large model framework to generate geological lithological material textures, adjust the training parameters according to the generated texture results, and re-execute step S300 until the preset texture result requirements are met. Use the diffusion model LDM large model framework and the trained weight matrix to create the interface API for generating geological lithological material textures.
[0113] Step S300 specifically includes the following steps:
[0114] Step S301, Training Preparation: Load the LDM large model framework and configure the training parameters;
[0115] Step S302, Start Training: Check the parameter update status of the LoCon model by the average loss rate of each training round. Adjust the training parameters according to the convergence status reflected by the parameters of the LoCon model, and retrain until the convergence speed meets the requirements.
[0116] Furthermore, in step S100, the pre-trained model uses a variational autoencoder (VAE), uses Clip as the text decoder, and uses cosion with restart as the scheduler.
[0117] Further, in step S100, after adding the low-rank matrix of LoCon to the pre-trained weight matrix of U-Net in the pre-trained model, the original weights are frozen.
[0118] Furthermore, in step S100, all components are loaded using the pipe encapsulated in the pipe.
[0119] Further, step S100 specifically integrates the convolutional layer into the convolutional weights of the residual block of the U-Net model in the following manner:
[0120] Parameters are extracted from the convolutional layers extended from the LoCon model to obtain the convolutional kernel parameter matrix WLoCon, and the weight matrix WU−Net of the target convolutional layer in the residual block of the U-Net model is obtained at the same time.
[0121] Add WLoCon to WU−Net element by element, that is, the new convolutional weight matrix Wnew=WU−Net+WLoCon, to achieve the integration of the convolutional layer;
[0122] Next, the low-rank matrix MLoCon of the LoCon model is added to the pre-trained weight matrix Wpretrained in the U-Net model by matrix addition, i.e., Wfinal=Wpretrained+MLoCon, to obtain the initial weight matrix, and then the geological lithology material texture model is built.
[0123] Furthermore, in step S200, images are collected from the network under the guidance of geological personnel, or sample photos are provided, and after texturization processing, the images are made to ensure that they reflect the target features.
[0124] Furthermore, in step S200, an image editing tool is used to perform pure texturing processing on the image. The processed image does not contain spatial orientation, perspective, lighting, or object outlines of three-dimensional information, but only reflects the color and features of the material surface.
[0125] Furthermore, in step S200, the samples are labeled with features using text descriptions or keywords, and a labeling tool plugin is used to generate a labeling file and corresponding icon name labels to generate a label set.
[0126] Furthermore, the feature annotation includes trigger words and additional descriptive words, with the trigger words being placed at the beginning of the label set as the main identifier for training.
[0127] Furthermore, the initial tag set is cleaned and corrected by deleting feature tags related to the target ontology in the training set; and the order of the remaining tags, except for the trigger words, is randomly shuffled.
[0128] Furthermore, the training set is reviewed and adjusted in batches, and erroneous labels and labels irrelevant to the target in the training set are deleted in batches; for images with common characteristics, descriptive labels of common characteristics are added in batches.
[0129] Furthermore, for individual image elements that the labeling tool fails to recognize, corresponding labels are added to the individual images through manual comparative analysis.
[0130] Furthermore, in step S200, data augmentation specifically includes: repeatedly copying a single feature sample in the training set, feature classification, multi-dimensional changes to the sample image, and shuffling a random seed.
[0131] Furthermore, feature classification includes: performing multi-bucket classification on images in the training set, cropping the texture of images within a preset range, and grouping them into the same bucket, with each bucket representing a set of images of a certain resolution.
[0132] Furthermore, the multidimensional variations of the samples include: performing color space adjustment operations and geometric transformation operations on the sample images. The color space adjustment includes brightness transformation, contrast adjustment, and hue transformation. The geometric transformation operations include horizontal flipping and vertical flipping to simulate the color characteristics and symmetry variations of the samples under different depositional environments.
[0133] Furthermore, in step S301, the training parameters set include the total number of steps, U-Net learning rate, TextEncoder learning rate, network dimension, and optimizer. After freezing the original pre-trained weights, the number of training epochs is set.
[0134] Further, in step S302, the convergence status of the LoCon model is observed by checking the average loss rate of each training epoch. If the convergence is too fast, the learning rate is reduced; if the convergence is too slow, the step size parameter is increased. The convergence speed is judged by comparing the number of training epochs at which convergence is achieved with a preset training epoch threshold. If the number of training epochs at which convergence is achieved is higher than the upper limit of the preset training epoch threshold range, the convergence speed is considered too slow; if the number of training epochs at which convergence is achieved is lower than the lower limit of the preset training epoch threshold range, the convergence speed is considered too fast.
[0135] Further, in step S400, the trained weights are output, the diffusion model LDM large model framework is loaded to generate geological lithological material textures, and the training parameters are adjusted based on the generated texture results, specifically including:
[0136] If the number of times a geological lithological texture model is generated exceeds a preset number, the similarity of the generated geological lithological texture effect is lower than the first preset threshold, or an irrelevant background is generated, then overfitting occurs, and the learning rate is reduced.
[0137] If the geological lithological texture model generates a geological lithological texture effect that does not contain some features or has a similarity to some features that is lower than the second preset threshold, then the number of repetitions of a single feature sample is increased.
[0138] If the generated geological lithology texture model produces a texture effect that has no texture or has a similarity of 0 with the expected texture, then underfitting occurs. In this case, increase the total number of steps and the learning rate, increase the Rank parameter in the LoCon model, decrease the Alpha parameter in the LoCon model, and increase the number of repetitions for a single feature sample.
[0139] Furthermore, step S100 specifically includes the following steps:
[0140] 1. Model initialization: Load the initial weight matrix `W_final` of the U-Net model, which contains the low-rank matrix of the LoCon model, into the residual block convolutional layer of the U-Net model as the initial values of the model parameters; freeze the parameters of the `W_pretrained` part (i.e. the original weights of U-Net), and only allow the LoCon parameters corresponding to `M_LoCon` (such as the `down` and `up` layers of the low-rank matrix) to participate in gradient updates.
[0141] 2. Forward propagation computation: Input training set sample images are converted into latent vectors through the encoding process of the LDM framework (VAE encoder); the latent vectors enter the U-Net model, and denoising is performed based on the initial weight matrix to generate predicted latent vectors; the predicted vectors are restored to the generated images by the VAE decoder, and the loss (such as L1 loss, perceptual loss) is calculated with the sample images.
[0142] 3. Backpropagation optimization: During backpropagation of the loss function, only the gradient of the LoCon parameter (`M_LoCon`) is calculated, and the gradient of the original weights of U-Net is set to zero. The optimizer (such as AdamW) updates the low-rank matrix parameters of LoCon according to the gradient, so that `M_LoCon` gradually fits the geological lithology characteristics, while `W_pretrained` remains unchanged.
[0143] 4. Iterative update logic: In each round of training, the `W_pretrained` part of the initial weight matrix is always used as the baseline, and the `M_LoCon` part is continuously adjusted through training; when the training reaches the preset number of rounds or the loss converges, the final `M_LoCon` parameters are combined with the fixed `W_pretrained` to form the complete weights after training.
[0144] Example 3
[0145] The present invention will now be described in detail with reference to a specific implementation scheme. For details not covered herein, please refer to Embodiment 2.
[0146] This invention provides a weight training method for generating geological lithological material texture models based on LoCon, comprising the following steps:
[0147] 1) such as Figure 2 As shown, a geological lithological material texture model is created.
[0148] Create the LoCon class;
[0149] Instantiate the LoCon class and generate a weight matrix;
[0150] Obtain the U-Net model and determine the target layer to mount;
[0151] Attach LoCon to the U-Net model;
[0152] Configure the training environment;
[0153] like Figure 3 This is the implementation code for the UNet core structure.
[0154] At this point, the geological lithology texture model has been created, and the next step is to wait for the training set images to be loaded for training.
[0155] 2) Perform preprocessing of training sample images
[0156] By repeatedly copying images with few samples and combining them with random spatial color and geometric transformations, a more complete training set is formed, such as... Figures 5-9 They are respectively Figure 4 Images obtained by processing the original image in different ways.
[0157] 3) Training the model to generate geological lithological material textures.
[0158] like Figure 10This is the training log for this embodiment. Epoch is the round number, Train_loss is the loss rate, Train_acc is the accuracy rate, and test_loss and test_cc are generated by first creating a graph using the weights of the current round, and then comparing it with the graph generated using the corresponding prompt words from the training set samples to check the loss rate and accuracy.
[0159] After each training round, the best-performing model is saved, and a backpropagation function is performed to calculate the new gradient. The optimizer is responsible for updating the parameters of the weighted model. After multiple rounds of parameter updates, the generated image features will become increasingly closer to the ideal, and will resemble the image features in the training set more and more. Of course, during this process, parameters such as the learning rate may need to be adjusted to help the gradient descent and converge faster, ultimately achieving the highest accuracy, which means the model's precision is the highest.
[0160] 4) Generate geological lithological material textures
[0161] The model was saved multiple times during training. Based on experience, the model in the last round isn't necessarily the best. It might be the model from three files prior to the last one. This .pth file is loaded, the prompt word is input, the inference function (texture generation program) is executed, a noise map is generated in the latent space, and denoising is performed based on the prompt word conditions, finally generating the required image, such as... Figure 11 As shown, the vertical axis represents the effect of the model saved in different batches, and the horizontal axis represents the scaling factor after the LoCon model is attached to the LDM basic model framework. The higher the value, the more obvious the effect of LoCon.
[0162] The above description is merely an optional embodiment of the present invention and does not limit the patent scope of the present invention. All equivalent structural transformations made using the contents of the present invention's specification and drawings under the inventive concept of the present invention, or direct / indirect applications in other related technical fields, are included within the patent protection scope of the present invention.
Claims
1. A weight training method for generating a geological lithology material texture model based on LoCon, characterized in that, Comprising the following steps: Step S100, using a latent diffusion model LDM as a large model framework for weight training and texture generation of a geological lithology material texture model, loading a pre-trained weight matrix using the denoising core of a U-Net model, expanding a convolution layer in a LoCon model, and integrating the convolution layer into the convolution weight of the residual block of the U-Net model, adding the low-rank matrix of the LoCon model to the pre-trained weight matrix in the U-Net model, and building a geological lithology material texture model; Step S200, obtaining a sample image of geological lithology material texture, performing texturing, feature labeling and data enhancement on the sample image to form a training set; Step S300, inputting the training set, training the generated geological lithology material texture model according to the pre-trained weight matrix added with the low-rank matrix of the LoCon model, until the preset training requirement is met; Step S400, outputting the trained weight, loading the diffusion model LDM large model framework to generate geological lithology material texture, adjusting the training parameters according to the generated result, and re-executing step S300 until the preset texture result requirement is met, and making an interface api between the diffusion model LDM large model framework and the trained weight matrix for geological lithology material texture generation; Wherein, step S300, specifically comprising the following steps: Step S301, training preparation: loading the diffusion model LDM large model framework and configuring the training parameters; Step S302, start training: check the parameter update state of the LoCon model through the average loss rate of each training round, adjust the training parameters according to the convergence state reflected by the parameters of the LoCon model, and retrain until the convergence speed meets the requirements. 2.The weight training method for generating a geological lithology material texture model based on LoCon according to claim 1, wherein, In step S100, the pre-training model uses a variational autoencoder VAE, uses Clip as a text decoder, and uses cosion with restart as a scheduler. 3.The weight training method for generating a geological lithology material texture model based on LoCon according to claim 1, wherein, In step S100, after adding the low-rank matrix of LoCon to the pre-trained weight matrix of U-Net in the pre-training model, the original weight is frozen.
4. The weight training method for generating a geological lithology material texture model based on LoCon according to claim 1, characterized in that, In step S100, all components are loaded using a pipe package. 5.The weight training method for generating a geological lithology material texture model based on LoCon according to claim 1, wherein, In step S200, pictures are collected from the network or sample photos are provided under the guidance of geologists, and after texturing, the images ensure that the target features are reflected.
6. The weight training method for generating a geological lithology material texture model based on LoCon according to claim 5, characterized in that, In step S200, the image is processed using a picture editing tool to obtain a pure texture, and the processed image does not contain three-dimensional information such as spatial orientation, perspective, light and shadow, or object contour, but only reflects the color and features of the material surface.
7. The weight training method for generating a geological lithology material texture model based on LoCon according to claim 1, wherein, In step S200, the sample is labeled by text description or keywords, and a labeling tool plug-in is used to generate a labeling file and corresponding icon name label, and a label set is generated.
8. The weight training method for generating a geological lithology material texture model based on LoCon according to claim 7, characterized in that, The feature label includes a trigger word and an additional description word, and the trigger word is placed at the beginning of the label set as the main body of the training.
9. The weight training method for generating a geological lithology material texture model based on LoCon according to claim 8, characterized in that, The initial label set is cleaned and corrected, the feature labels related to the target ontology in the training set are deleted, and the remaining labels except the trigger word are randomly shuffled.
10. The weight training method for generating a geological lithology material texture model based on LoCon according to claim 7, characterized in that, The training set is audited and batch adjusted, and the error labels and labels irrelevant to the target in the training set are batch deleted; for images with common characteristics, the descriptive labels of the common characteristics are batch added.
11. The weight training method for generating a geological lithology material texture model based on LoCon according to claim 7, characterized in that, For single image elements not recognized by the labeling tool, the corresponding labels are supplemented for the single image through manual comparative analysis.
12. The weight training method for generating a geological lithology material texture model based on LoCon according to claim 1, wherein, In step S200, data enhancement specifically includes: repeating and duplicating individual feature samples in the training set, feature classification, multi-dimensional changes of sample images, and random seed shuffling.
13. The weight training method for generating a geological lithology material texture model based on LoCon according to claim 12, characterized in that, Feature classification includes: multi-bucket classification of pictures in the training set, texture cropping of images with a proportion within a preset range, and grouping into the same bucket, each bucket representing a collection of images of one resolution.
14. The weight training method for generating a geological lithology material texture model based on LoCon according to claim 12, characterized in that, The multi-dimensional changes of the sample include: performing color space adjustment operation and geometric transformation operation on the sample image, the color space adjustment includes brightness transformation, contrast adjustment and hue change; the geometric transformation operation includes horizontal flip and vertical flip, to simulate the color characteristics and symmetric variants of the sample under different deposition environments.
15. The weight training method for generating a geological lithology material texture model based on LoCon according to claim 1, wherein, In step S301, the training parameters set include the total number of steps, the learning rate of U-Net, the learning rate of Text Encoder, the network dimension, the optimizer, and the training number of epochs after freezing the original pre-trained weights.
16. The weight training method for generating a geological lithology material texture model based on LoCon according to claim 1, wherein, In step S302, the convergence state of the parameter update state of the LoCon model is viewed through the average loss rate of each training round, if the convergence is too fast, the learning rate is reduced, if the convergence speed is too slow, the step parameter is increased, the convergence speed is too fast or too slow, which is judged by comparing the training round when the convergence is reached with the preset training round threshold, if the training round when the convergence is reached is higher than the upper limit of the preset training round threshold range, it is considered that the convergence speed is too slow, if the training round when the convergence is reached is lower than the lower limit of the preset training round threshold range, it is considered that the convergence speed is too fast.
17. The weight training method for generating a geological lithology material texture model based on LoCon according to claim 1, wherein, In step S400, the trained weights are output, the diffusion model LDM large model framework is loaded to generate geological lithology material texture, and the training parameters are adjusted according to the generated texture result, specifically including: If the geological lithology material texture generated by the generated geological lithology material texture model exceeds the preset number of times, the similarity of the generated geological lithology material texture effect is lower than the first preset threshold, and irrelevant background is generated, overfitting is generated, and the learning rate is reduced; If the geological lithology material texture generated by the generated geological lithology material texture model does not contain part of the feature or the similarity with part of the feature is lower than the second preset threshold, the number of repetitions of the single feature sample is increased; If the generated geological lithology material texture generated by the generated geological lithology material texture model has no texture or the similarity with the expected texture is 0, underfitting is generated, the total number of steps and the learning rate are increased, the parameter Rank in the LoCon model is increased, the parameter Alpha in the LoCon model is reduced, and the number of repetitions of the single feature sample is increased. 18.A weight training system for generating a geological lithology material texture model based on LoCon, characterized in that, The processor can execute a computer program, and the computer program can implement the weight training method of the LoCon-based geological lithology material texture model generation method according to any one of claims 1-17.
Citation Information
Patent Citations
Intelligent substation data enhancement and anomaly identification method
CN117788976A
Content synthesis using generative Artificial Intelligence model
US12271978B1