Material image database generation method and system based on large language model
Through the material image database generation method based on the large language model, cloud servers and local hosts work together, and combined with LoRA and TI algorithms to fine-tune the literary image model, the consistency and accuracy problems in the generation of material image databases are solved, and high-quality material images are automatically generated and edited.
Patent Information
- Application Number
- CN202510568880.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-15
AI Technical Summary
In the prior art, the material image database generation has problems such as poor manual label consistency, high cost and poor accuracy of the label generation algorithm, making it difficult to generate high-quality and consistent material images.
Using a method based on a large language model, a cloud server and local host work together, a structured material description is generated using a large language model, and a material image database is automatically generated, including the material image description design and image style derivation stage, using LoRA and TI algorithms for model fine-tuning and image style retrieval.
It realizes high-quality and consistent material image generation, supports material attribute editing, has high accuracy in generation and description, and has good retention of the original material concept, and is suitable for large-scale data annotation scenarios.
Smart Images

Figure CN120492649A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of material image database generation, and in particular to a material image database generation method and system based on a large language model. Background Art
[0002] The training (and fine-tuning) process of the text-based graph model uses a database of texture images as training data. Achieving fine-grained control over the generation of text-based graph models and generating the high-quality images required by users requires a high-quality texture database as the training data foundation.
[0003] Existing technology material image database generation process:
[0004] 1. Manual Annotation: Currently, a considerable portion of datasets are still manually annotated by researchers. However, manual annotation is subject to individual differences in researchers' subjective experience, resulting in poor consistency. Furthermore, manual annotation is costly and unsuitable for large-scale data annotation scenarios.
[0005] 2. Label Generation Algorithms: Common label generation algorithms include Deepbooru and BLIP, both of which can be used to generate material image annotations. However, these two label generation algorithms suffer from poor accuracy and difficulty in controlling the format of generated content. Summary of the Invention
[0006] In order to address the deficiencies of the prior art, the present invention provides a method and system for generating a material image database based on a large language model;
[0007] On the one hand, a method for generating a material image database based on a large language model is provided;
[0008] A method for generating a material image database based on a large language model includes:
[0009] (1) The local host uploads the target material image and the preset material image prompt words to the cloud server;
[0010] (2) The cloud server receives the target material image and the preset material image prompt words, and inputs the target material image and the preset material image prompt words into the large language model of the cloud server, and the large language model outputs a structured material description; the cloud server sends the structured material description to the local host;
[0011] (3) The local host inputs the placeholder corresponding to the target material image style and the structured material description into the trained text graph model, and the trained text graph model outputs a derived material image;
[0012] (4) Replace the target material image and repeat (1)-(3) to obtain several derived material images. Summarize all derived material images and the structured material descriptions corresponding to the images to obtain a material image database.
[0013] On the other hand, a material image database generation system based on a large language model is provided;
[0014] A material image database generation system based on a large language model, including: a local host and a cloud server;
[0015] The local host uploads the target material image and the preset material image prompt words to the cloud server;
[0016] The cloud server receives the target material image and the preset material image prompt word, inputs the target material image and the preset material image prompt word into the cloud server's large language model, and the large language model outputs a structured material description; the cloud server sends the structured material description to the local host;
[0017] The local host inputs the placeholder corresponding to the target material image style and the structured material description into the trained text-based graph model, and the trained text-based graph model outputs a derived material image;
[0018] The target material image is replaced to obtain several derivative material images, and all the derivative material images and the structured material descriptions corresponding to the images are summarized to obtain a material image database.
[0019] The above technical solution has the following advantages or beneficial effects:
[0020] Material descriptions can be automatically generated; the generated material descriptions have consistent style and high accuracy, and conform to the physical characteristics of the material image; users can edit the physical properties of the material image by modifying part of the material description; the demand for target material samples required to generate derived materials is small, and the concept of the original material is well retained. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0022] Figure 1 This is a flow chart of the technology for generating a material image database based on a large language model according to the present invention.
[0023] Figure 2 A flow chart for designing material image description in the method of the present invention.
[0024] Figure 3This is a flow chart of material image style derivation in the method of the present invention;
[0025] Figure 3 (a) in the figure shows the results generated by the untrained Wensheng graph model;
[0026] Figure 3 (b) shows the generation results of the LoRA-trained text graph model in the two-dimensional material domain;
[0027] Figure 3 (c) in the figure shows the generation results of the textual graph model trained with Textural Inversion in the target material derivative field. DETAILED DESCRIPTION
[0028] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.
[0029] Example 1
[0030] This embodiment provides a method for generating a material image database based on a large language model;
[0031] like Figure 1 As shown, the method for generating a material image database based on a large language model includes:
[0032] S101: The local host uploads the target material image and the preset material image prompt words to the cloud server;
[0033] S102: The cloud server receives a target material image and a preset material image prompt word, inputs the target material image and the preset material image prompt word into a large language model of the cloud server, and the large language model outputs a structured material description; the cloud server sends the structured material description to the local host;
[0034] S103: The local host inputs the placeholder corresponding to the target material image style and the structured material description into the trained text graph model, and the trained text graph model outputs a derived material image;
[0035] S104: Replace the target material image and repeat S101-S103 to obtain several derived material images. Aggregate all derived material images and the structured material descriptions corresponding to the images to obtain a material image database.
[0036] Furthermore, the target material image includes: several material images, such as: wood (including wooden floors, logs, and wooden utensils), ceramics (including shiny ceramics and rough ceramics), fabrics (including knitted fabrics, denim, and wool fabrics), metals (including polished metals, matte metals, and brushed metals), glass (including transparent, translucent, colorless, or transparent glass), and stone (including marble and granite).
[0037] Furthermore, the preset material image prompt words include:
[0038] The prompt word is clear and specific. While asking the question itself, it should also impose a series of restrictions on the answer. The prompt word specifies the steps to complete the task, breaking down complex tasks into multiple steps and guiding the large language model to complete them in sequence. The prompt word controls the output format, requiring the large language model to output in the target format.
[0039] Furthermore, the preset material image prompt words include:
[0040] Set up the role of the large language model;
[0041] The task of setting up a large language model is to provide a structured material description of the image;
[0042] Set the application areas of large language models;
[0043] Set the format of the large language model output text.
[0044] The format of the output text, including: material color, material surface texture, whether the material is isotropic or anisotropic, the name of the material's components, and the material's roughness.
[0045] Furthermore, in S102, the cloud server receives a target material image and preset material image prompt words, and the cloud server inputs the target material image and the preset material image prompt words into a large language model of the cloud server, and the large language model outputs a structured material description; wherein the large language model is implemented using a GHAT GPT model or a DeepSeek model.
[0046] Furthermore, the cloud server is provided with an interface and a key, and the interface and the key are used to monitor and process batch access programs of the local host.
[0047] The structured material description includes: material color, material surface texture, whether the material is isotropic or anisotropic, material component names, and material roughness.
[0048] Furthermore, in S103 : the local host inputs the placeholder corresponding to the target material image style and the structured material description into the trained text graph model, and the trained text graph model outputs a newly generated material image;
[0049] The target material image style includes: the sub-category of the material;
[0050] If the target material image is a wood image, the material image styles include: wooden floor, log, and wooden utensils;
[0051] If the target material image is a ceramic image, the material image styles include: glossy ceramic and rough ceramic;
[0052] If the target material image is a fabric image, the material image styles include: knitted fabric, denim fabric, and wool fabric;
[0053] If the target material image is a metal image, the material image style includes: polished metal, matte metal, brushed metal;
[0054] If the target material image is a glass image, the material style image includes: transparent glass, translucent glass;
[0055] If the target material image is a stone image, the material style image includes: marble or granite.
[0056] Among them, the placeholder corresponding to the target material image style refers to: pseudo-word, which is a text encoding vector.
[0057] Among them, the placeholder corresponding to the target material image style is a word, and the concept corresponding to the word represents a material image style.
[0058] For example, for a target material being a wooden table surface, the placeholder may be defined as "WOODSURF1".
[0059] Among them, the cultural graph model is implemented using Stable Diffusion 1.5.
[0060] The trained Wensheng graph model includes: two rounds of training;
[0061] First round of training: Construct a first training set, which is a structured text description of a known two-dimensional material image (e.g., a texture image of a material); input the first training set into the Vincent graph model to train the model. When the model's loss function value no longer decreases, or the number of iterations exceeds a set number, stop training, and obtain the Vincent graph model after the first round of training;
[0062] Second round of training: Construct a second training set, which consists of structured text descriptions of known material images and placeholders corresponding to image styles; input the second training set into the Wenshengtu model after the first round of training, using the structured text descriptions and placeholders corresponding to image styles as the input values of the model, and the known material images as the output values of the model. When the model loss function value no longer decreases, or the number of iterations exceeds the set number, stop training, and obtain the Wenshengtu model after the second round of training;
[0063] The Wensheng graph model after the second round of training is used as the final Wensheng graph model after training.
[0064] Furthermore, in the first round of training, the LoRA (Low-Rank Adaptation) algorithm is used to fine-tune the parameters of the text graph model.
[0065] Furthermore, the LoRA algorithm is used to fine-tune the parameters of the cultural graph model, specifically including:
[0066] LoRA parameter initialization;
[0067] Input image-text pair and optimize parameters by gradient descent;
[0068] When the set training deployment conditions are reached, stop training.
[0069] Furthermore, in the second round of training, a textual inversion (TI) algorithm is used to determine the material domain corresponding to the target material in the latent space of the text graph model.
[0070] Furthermore, the text inversion TI algorithm is used to determine the material domain corresponding to the target material in the latent space of the text graph model, including:
[0071] (1) Input a set of 3 to 5 target material images and a placeholder (pseudo-word), which is a custom word used to describe the style of the target material image;
[0072] (2) Optimizing the latent vector in the text encoder of the text-based graph model corresponding to the placeholder through multiple rounds of training, so that the latent vector represents the style of the target material image in the latent space of the text encoder;
[0073] The optimized implementation is as follows:
[0074]
[0075] Among them, v *is the latent vector of the image style in the latent space of the Wensheng graph model, z is the latent vector in the latent space of the Wensheng graph model, ε is the image encoder, x is the target material sample image, y is the prompt word template corresponding to the placeholder embedded in the image style, ∈ is the noise added in the process of denoising the Wensheng graph model, t is the time step in the Wensheng graph model, ∈ θ is the noise generator of the Wensheng graph model, c θ A text encoder. For expectations, is a standard normal distribution.
[0076] (3) After reaching the set number of training steps, the training stops.
[0077] This invention targets scenarios where existing materials need to be modified and derived, ensuring that the style of the new material remains consistent with the original while completing detailed editing of the material. This method utilizes a large language model-based text-graph model material derivation technique. This method uses the large language model to generate material descriptions and uses the text-graph model to perform image style retrieval, ultimately enabling material editing and derivation.
[0078] The material image database generation technology based on the large language model is divided into two stages: the material image description design stage and the material image style derivation stage.
[0079] In the material image description design stage, the first thing to do is to design a structured description of the material image. The characteristics of the material include color, surface texture (such as concave and convex patterns, brushed patterns, textile patterns), transparency, etc. These characteristics define the physical properties of a material and determine the optical appearance of the material. In the application of text image, not all descriptions of material characteristics have a positive impact on improving the quality of material image generation. Therefore, in the material image description design stage, it is necessary to design a set of material structure descriptions that are suitable for the input of the text image model. Traditional material descriptions are manually annotated by researchers. This manual material annotation is largely affected by the subjective experience of researchers and has poor consistency. Using a multimodal large language model batch processing guided by prompt words can generate high-quality material descriptions with excellent consistency.
[0080] In the material image style derivation stage, the present invention uses the Vincent graph model to generate an image style derived material with the same style as the target material sample. The Vincent graph model needs to be adjusted to achieve the image style derivation of the material, and the main steps of the adjustment include: using material domain samples to fine-tune the Vincent graph model, and initially limiting the output of the Vincent graph model to the two-dimensional material domain. After limiting the output of the Vincent graph model to the two-dimensional material domain, the present invention uses a small number of target material samples to retrieve the position of the image style corresponding to the target material sample in the latent space, and uses a placeholder to represent this image style. Using this placeholder in combination with a structured material description, the present invention can realize the batch generation of derived target materials for a target material representing an image style, and supports users to obtain customized derived materials by modifying the structured material description.
[0081] The present invention provides a fine-grained image generation technology targeting material images, which ensures the consistency of material style as much as possible while taking into account the flexibility of material editing.
[0082] The specific process of designing a structured description of a material image is as follows: In the field of physically based rendering (PBR), a set of standardized material properties are usually used to describe the characteristics of a material. The surface of an object seen by the human eye is determined by the light propagating from the surface of the object to the human eye (including reflection, refraction, self-illumination, etc.). To simulate the propagation of light on objects in the physical world, PBR defines a bidirectional reflectance distribution function (BRDF). For example, in the classic Cook-Torrance model, the reflection of light on the surface of a material can be defined as:
[0083] f r =k d f lambert +k s f cook-torrance
[0084] Among them, k d is the diffuse reflectance, which represents the proportion of diffusely reflected light in the incident light; k s Is the specular reflection coefficient, which represents the proportion of light that is specularly reflected in the incident light; the reflection model of the diffuse reflection part
[0085] Reflection model of the specular reflection part:
[0086]
[0087] Among them, c is the base color of the material, D is the normal distribution function of the material, F is the Fresnel equation that describes the proportion of light reflected from the surface under different incident angles, and G is the geometric function that describes the light obstruction caused by roughness.
[0088] From the Cook-Torrance model's simulation of light reflection on a material surface, we can see that the color, roughness, and geometric shape of a material all affect its optical effect. Based on the understanding of material properties in rendering, this paper designs a structured description of a material image. This structured description of a material image is a set of material properties composed of independent members. The set members include:
[0089] Material color, material surface texture (such as polishing, brushing, carving, etc.), whether the material is isotropic or anisotropic, the name of the main components that make up the material, and the roughness of the material.
[0090] On the basis of basic reflection, the present invention also takes into account the refraction and projection of light in the material and adds properties such as transparency.
[0091] The structured material image is based on knowledge from physically based rendering. Material color describes the color of the material, while material surface texture describes the distribution of the material's normals. The names of the material's main components and whether the material is isotropic or anisotropic describe the proportion of light reflected from the surface at different incident angles. Material roughness describes the light obstruction caused by the roughness of the material. This structured material description has a sound theoretical basis and effectively preserves the information in the material image.
[0092] Large language model prompt word engineering design: Large language model prompt word engineering design serves the design of structured descriptions of material images. For the task of generating material descriptions, the traditional solution is for researchers to manually annotate material images. This manual annotation description generation method has many drawbacks: First, there are large individual differences in researchers' professional experience and subjective cognition. The material description library generated by multiple researchers writing material descriptions in parallel often has consistency defects, which is not conducive to downstream applications (such as the model's learning of material image style during fine-tuning training of text-based image models). Second, the time and economic costs of hiring researchers to write material descriptions are high, which is difficult to implement in the task of generating large-scale material description libraries.
[0093] In recent years, multimodal large language models have been widely used in various fields due to their powerful semantic recognition and logical analysis capabilities. As the standard input driving large language models, prompt words have attracted extensive attention from researchers. Excellent prompt words can better guide large multimodal language models to complete their tasks, thus giving rise to a new research area: prompt word engineering.
[0094] To generate structured descriptions of material images, this method designs a set of prompts based on the large language model prompts project, with clear command constraints. These prompts require the large language model to act as a material describer, generating a material description for the input material image sample in the format of a structured material image description.
[0095] Large models require prompts to achieve tasks. Prompts with clear task constraints and complete constraints can maximize the performance of large models. During the large language model prompt engineering design step, a prompt for the material description task was designed. Its specific content includes the following parts:
[0096] 1. Describe the role of the large model in the conversation. The prompt word first needs to inform the large model of its general task domain. In this method, the prompt word first informs the large model of its role in the conversation: material describer. Then, the prompt word needs to inform the large model of the detailed content of this task. In this method, the prompt word informs the large model that the material describer's task is to provide material descriptions for the material images uploaded in the conversation, and indicates that the application domain of this material description is the text-based image model.
[0097] Second, fine-grained constraints on the task. In the first part, the prompt word informed the large model of its role and the corresponding task. In this part, the prompt word also needs to inform the large model of the key points to pay attention to in order to complete the task. In this method, the prompt word tells the large model to focus only on the surface of the subject in the material image and not on any other details in the image (such as the subject's shape, background, etc.).
[0098] 3. Constraints on the output format. The content output by the large model in the conversation is the material description of the material by the material describer, which in this method is the content of the structured description design of the material image. In this method, the prompt word will inform the large model that the conversation response must include only the material color, the material surface texture (such as polished, brushed, engraved, etc.), whether the material is isotropic or anisotropic, the name of the main component of the material, the material roughness, and the material transparency. The output content is separated by commas, and the end symbol is not retained.
[0099] The configuration process for a large language model in the cloud is as follows: Because large language models require high computing power from computer hardware, deploying them locally is difficult in practice. This paper adopts a method for configuring and invoking a large language model in the cloud. First, the large language model is deployed on a cloud server, and an interface and key are set up to monitor and process access from local batch programs. Then, based on task requirements, the large language model's knowledge base and plug-ins are adjusted to ensure that it has the basic knowledge and functionality to generate high-quality material descriptions.
[0100] The local batch processing model is configured as follows: In this method, all material samples are stored on the local host, and multimodal batch processing of a large language model is achieved through a batch processing program running on the local host. The local batch processing program first uploads the locally stored material sample images to a cloud server, then sends a query request to the server and polls the cloud server for a response. The query request includes a prompt word, and the server response contains a structured description of the material image.
[0101] like Figure 2 As shown, the local host uploads material sample images from the material sample database to the cloud server in batches. After the upload is complete, the local batch processing model sends a query request to the cloud server, prompting the user to establish a material description session. The large cloud model then enters the session processing state (i.e., generating a structured material description). This method incorporates a polling program within the local batch processing model to periodically query the status of the large cloud model. When the large cloud model enters the processing completed state (i.e., structured material description generation is complete), the local batch processing model receives the session response and obtains the structured material description.
[0102] like Figure 3 As shown, the material image style derivation part is mainly the adjustment of the cultural graph model. The specific process of material domain constraint generation of the cultural graph model is as follows: Based on the multimodal batch processing application of the large language model, this method can generate a corresponding structured material description dataset for a large-scale material image dataset. This method uses the LoRA (Low-Rank Adaptation) method to fine-tune the cultural graph model. The LoRA method uses the low-rank characteristics of the model parameter matrix to simulate full-parameter fine-tuning by embedding a small-scale bypass matrix inside the model. Compared with the full-parameter fine-tuning method, the computational cost of this method is significantly reduced, and there is no obvious impact on the loss of model performance. After fine-tuning on the structured material description dataset, the generation domain of the cultural graph model is constrained to the two-dimensional material domain, realizing two-dimensional material domain generation.
[0103] Generating material domain constraints in the text graph model. The original text graph model is similar to the large language model. The generation domain is unconstrained and includes many things in all fields. This method focuses on the field of material generation in the field of computer graphics generation. It uses the LoRA method to fine-tune the parameters of the bypass matrix of the generation model to constrain the generation domain of the text graph model. The material description used for fine-tuning is provided by the large language model in the material image description design part. This method provides a large number of high-quality structured material descriptions. The results before and after fine-tuning are compared. Figure 3As shown in (a) and (b) of Figure 1, for the same text-based image cue, before fine-tuning, the generated images by the text-based image model exhibit significant perspective randomness, and the generated images do not represent texture images. After fine-tuning, the generated domain of the text-based image model is constrained to the 2D texture domain, generating texture images. Figure 3 (a) in the figure shows the results generated by the untrained Wensheng graph model; Figure 3 (b) shows the generation results of the LoRA-trained text graph model in the two-dimensional material domain; Figure 3 (c) in the figure shows the generation results of the textual graph model trained with Textural Inversion in the target material derivative field.
[0104] The specific process of image style retrieval using the text-based graph model is as follows: The LoRA method is used to fine-tune the text-based graph model, limiting its generative domain to the two-dimensional material domain. To achieve material image style derivation, this method requires obtaining the material domain corresponding to the target material in the latent space of the text-based graph model. This method uses the TI (Textural Inversion) method to determine the material domain corresponding to the target material in the latent space of the text-based graph model. The TI method freezes the parameters of the text encoder and the inversion network in the text-based graph model, and only optimizes the latent vector of the placeholder corresponding to the image style. The optimization is implemented as follows:
[0105]
[0106] Among them, v * is the latent vector of the image style in the latent space of the Wensheng graph model, z is the latent vector in the latent space of the Wensheng graph model, ε is the image encoder, x is the target material sample image, y is the prompt word template corresponding to the placeholder embedded in the image style, ∈ is the noise added in the process of denoising the Wensheng graph model, t is the time step in the Wensheng graph model, ∈ θ is the noise generator of the Wensheng graph model, c θ A text encoder. For expectations, is a standard normal distribution.
[0107] The TI method is used to further retrieve the image style corresponding to the target material in the two-dimensional material domain. The TI method requires a small number of target material images (about 5). By reversely optimizing the latent vector of the image style corresponding to the target material in the latent space of the text graph model, the vector is bound to a custom placeholder. After training, the custom placeholder and natural language can be combined in the text graph model to realize the text graph model image style retrieval. The results are as follows: Figure 3 As shown in (c) in .
[0108] The process for generating a material image style derivation database based on structured descriptions is as follows: The text-based graph model possesses material image style derivation capabilities. The image style corresponding to a target material sample is retrieved and then invoked using a placeholder and natural language. The stochastic nature of the text-based graph model combined with the precise constraints of image style retrieval enables the model to derive material images with the same style as the target material sample. These material images can then share the same structured material description as the target material sample, constructing a target material library consisting of target-like material images. This method also allows for the construction of a customized derivative material library by customizing some members of the structured material description.
[0109] Generating a database of material image style derivatives based on structured descriptions. By combining material images generated by the text-based graph model image style retrieval with high-quality structured material descriptions generated by multimodal batch processing using a large language model, a database of material images derived from the target material can be constructed. Because this method utilizes structured material descriptions, users can freely edit the material properties within the descriptions, using customized material descriptions as prompts in the text-based graph model. This allows the generation of a customized database of material image style derivatives using methods from the text-based graph model image style retrieval.
[0110] In the material image description design phase, we design a structured, highly information-retaining material image description based on the physical properties of materials and the specific requirements for describing material attributes when generating text-based graph models. This material description can be used to build a high-quality structured database.
[0111] In the design of material image description, the large language model prompt word engineering is used to guide the large language model to generate structured material image descriptions with high information retention.
[0112] In the design of material image description, a batch processing program is designed to realize multimodal batch processing applications of large language models.
[0113] In the material image style derivation part, the material domain constraints of the text-based graph model are generated. The constraints are generated by training the text-based graph model with material domain samples under the consistency fine-tuning method.
[0114] In the material image style derivation part, the text-based graph model image style retrieval is performed. The image style retrieval is obtained by training the text-based graph model under latent space retrieval using target material samples.
[0115] In the material image style derivation part, a material image style derivation database is generated based on structured descriptions. The text-based graph model for image style retrieval can combine material descriptions and user needs to achieve material image style derivation.
[0116] The batch processing process is as follows: create a cloud-based multimodal large language model dialogue link, distribute material images and large language model prompts in batches, poll and wait for the cloud-based material description to be generated, and copy the material description of this batch.
[0117] This invention discloses a technology for generating a material image database based on a large language model, comprising a material image description design component and a material image style derivation component. The material image description design component includes: designing a structured material image description, engineering large language model prompt words, and applying the large language model's multimodal batch processing. The material image style derivation component includes: generating material domain constraints using a text-based graph model, searching for image styles using the text-based graph model, and generating a material image style derivation database based on structured descriptions. This technology features low sample requirements, high description accuracy, and good concept derivation retention.
[0118] Example 2
[0119] This embodiment provides a material image database generation system based on a large language model, including: a local host and a cloud server;
[0120] The local host uploads the target material image and the preset material image prompt words to the cloud server;
[0121] The cloud server receives the target material image and the preset material image prompt word, inputs the target material image and the preset material image prompt word into the cloud server's large language model, and the large language model outputs a structured material description; the cloud server sends the structured material description to the local host;
[0122] The local host inputs the placeholder corresponding to the target material image style and the structured material description into the trained text-based graph model, and the trained text-based graph model outputs a derived material image;
[0123] The target material image is replaced to obtain several derivative material images, and all the derivative material images and the structured material descriptions corresponding to the images are summarized to obtain a material image database.
[0124] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A method for generating a material image database based on a large language model, characterized in that: include: (1) The local host uploads the target material image and the preset material image prompt words to the cloud server; (2) The cloud server receives the target material image and the preset material image prompt words, and inputs the target material image and the preset material image prompt words into the large language model of the cloud server, and the large language model outputs a structured material description; the cloud server sends the structured material description to the local host; (3) The local host inputs the placeholder corresponding to the target material image style and the structured material description into the trained text graph model, and the trained text graph model outputs a derived material image; (4) Replace the target material image and repeat (1)-(3) to obtain several derived material images. Summarize all derived material images and the structured material descriptions corresponding to the images to obtain a material image database.
2. The method for generating a texture image database based on a large language model according to claim 1, wherein: The preset material image prompt words include: Set the role of the large language model; set the task of the large language model to provide structured material description for the image; set the application field of the large language model; set the format of the output text of the large language model.
3. The method for generating a texture image database based on a large language model according to claim 2, wherein: The format of the output text, including: material color, material surface texture, whether the material is isotropic or anisotropic, the name of the material's components, and the material's roughness.
4. The method for generating a texture image database based on a large language model according to claim 1, wherein: The cloud server is provided with an interface and a key, and the interface and the key are used to monitor and process batch access programs of the local host.
5. The method for generating a texture image database based on a large language model according to claim 1, wherein: The local host inputs the placeholder corresponding to the target material image style and the structured material description into the trained text-based graph model, and the trained text-based graph model outputs a newly generated material image; The target material image style includes: the sub-category of the material; If the target material image is a wood image, the material image styles include: wooden floor, log, and wooden utensils; If the target material image is a ceramic image, the material image styles include: glossy ceramic and rough ceramic; If the target material image is a fabric image, the material image styles include: knitted fabric, denim fabric, and wool fabric; If the target material image is a metal image, the material image style includes: polished metal, matte metal, brushed metal; If the target material image is a glass image, the material style image includes: transparent glass, translucent glass; If the target material image is a stone image, the material style image includes: marble or granite.
6. The method for generating a texture image database based on a large language model according to claim 1, wherein: The placeholder corresponding to the target material image style is a word, and the concept corresponding to the word represents a material image style.
7. The method for generating a texture image database based on a large language model according to claim 1, wherein: The trained Wensheng graph model includes: two rounds of training; The first round of training: Construct the first training set, which is a structured text description of a known two-dimensional material image; input the first training set into the text-graph model and train the model. When the loss function value of the model no longer decreases or the number of iterations exceeds the set number, stop training and obtain Wensheng graph model after the first round of training; Second round of training: Construct a second training set, which consists of structured text descriptions of known material images and placeholders corresponding to image styles; input the second training set into the Wenshengtu model after the first round of training, using the structured text descriptions and placeholders corresponding to image styles as the input values of the model, and the known material images as the output values of the model. When the model loss function value no longer decreases, or the number of iterations exceeds the set number, stop training, and obtain the Wenshengtu model after the second round of training; The Wensheng graph model after the second round of training is used as the final Wensheng graph model after training.
8. The method for generating a texture image database based on a large language model according to claim 7, wherein: In the first round of training, the LoRA algorithm is used to fine-tune the parameters of the text-graph model. The LoRA algorithm is used to fine-tune the parameters of the text-graph model, which specifically includes: initializing the LoRA parameters; inputting image-text pairs and optimizing the parameters through the gradient descent method; and stopping training after reaching the set training deployment conditions.
9. The method for generating a texture image database based on a large language model according to claim 7, wherein: The second round of training uses a text inversion algorithm to determine the material domain corresponding to the target material in the latent space of the Wensheng graph model. The use of the text inversion TI algorithm to determine the material domain corresponding to the target material in the latent space of the Wensheng graph model includes: (1) Input a set of target material images and a placeholder, wherein the placeholder is a custom word used to describe the style of the target material image; (2) Optimizing the latent vector in the text encoder of the text-based graph model corresponding to the placeholder through multiple rounds of training, so that the latent vector represents the style of the target material image in the latent space of the text encoder; The optimized implementation is as follows: Among them, v * is the latent vector of the image style in the latent space of the Wensheng graph model, z is the latent vector in the latent space of the Wensheng graph model, ε is the image encoder, x is the target material sample image, y is the prompt word template corresponding to the placeholder embedded in the image style, ∈ is the noise added in the process of denoising the Wensheng graph model, t is the time step in the Wensheng graph model, ∈ θ is the noise generator of the Wensheng graph model, c θ For text encoder; For expectations, is a standard normal distribution; (3) After reaching the set number of training steps, the training stops.
10. A material image database generation system based on a large language model, characterized in that: include: Local host and cloud server; The local host uploads the target material image and the preset material image prompt words to the cloud server; The cloud server receives the target material image and the preset material image prompt word, inputs the target material image and the preset material image prompt word into the cloud server's large language model, and the large language model outputs a structured material description; the cloud server sends the structured material description to the local host; The local host inputs the placeholder corresponding to the target material image style and the structured material description into the trained text-based graph model, and the trained text-based graph model outputs a derived material image; The target material image is replaced to obtain several derivative material images, and all the derivative material images and the structured material descriptions corresponding to the images are summarized to obtain a material image database.