Metallographic image generation method, device, equipment and medium
By combining the stable diffusion model and the LoRA model and using mask data to train the model, the problem of poor accuracy of existing metallographic image generation methods is solved, and more stable and high-quality metallographic image generation is achieved.
Patent Information
- Application Number
- CN202510166050.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-05-16
AI Technical Summary
The existing metallographic image generation methods rely on deep learning models, and have poor accuracy when processing microstructures, resulting in poor generation results.
A model combining a stable diffusion model and a LoRA model is adopted to fine-tune the cross attention layer weight parameters of the stable diffusion model through the LoRA model, and the model is trained with mask data to improve the image generation accuracy of the region of interest.
The stability and detail clarity of metallographic image generation are improved, the generation accuracy of the region of interest is enhanced, and the overall image generation effect is improved.
Smart Images

Figure CN120014093A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of image generation technology, and in particular, relates to a metallographic image generation method, device, equipment and medium. Background Art
[0002] Metallographic analysis is an important means to study the microstructure of metals and their alloys. Traditional metallographic image acquisition methods mainly rely on microscopic observation, prepare samples through manual sample preparation, polishing and etching, and then use optical microscopes or scanning electron microscopes (SEM) and other equipment to take metallographic images. The above methods can provide high-precision microstructural information, but the process is cumbersome and manual, time-consuming and difficult to efficiently process a large number of samples. At present, with the development of computer image processing technology, image generation methods based on image processing have emerged. The metallographic image generation method in the related art mainly relies on deep learning models such as generative adversarial networks, which often have poor accuracy when dealing with the fineness and complexity of the microstructure, which in turn leads to poor metallographic image generation. Summary of the invention
[0003] In view of the above-mentioned shortcomings of the prior art, the purpose of the present application is to provide a metallographic image generation method, device, equipment and medium to solve the above-mentioned problems.
[0004] The metallographic image generation method provided in this application includes:
[0005] Acquire a sample set of metallographic images, wherein the sample set includes a plurality of metallographic image data and corresponding text labels, wherein each of the plurality of metallographic image data includes a preprocessed metallographic image and mask data, and the mask data is used to mark a region of interest in the preprocessed metallographic image;
[0006] Training the first model based on the sample set to obtain a second model, wherein the first model is a model combining a stable diffusion model and a LoRA model, and the LoRA model is used to update the weight parameters of the cross attention layer of the stable diffusion model during the training process;
[0007] The text to be processed is input into the second model to obtain a first metallographic image corresponding to the text to be processed, wherein the text to be processed includes metallographic features.
[0008] In one embodiment of the present application, in the second model, the sampling time step is positively correlated with the regional weight, and the regional weight is the weight value assigned by the second model to different regions of the first metallographic image.
[0009] In one embodiment of the present application, the pre-processed metallographic image is determined by the following steps:
[0010] Collect original metallographic images;
[0011] Clipping the original metallographic image according to a preset size to obtain a clipped original metallographic image;
[0012] The cropped original metallographic image is subjected to image enhancement processing to obtain the preprocessed metallographic image.
[0013] In one embodiment of the present application, the region of interest in the preprocessed metallographic image is: a grain image region and an inclusion image region in the preprocessed metallographic image.
[0014] In an embodiment of the present application, the mask data is binary mask data.
[0015] In one embodiment of the present application, the text label of the metallographic image data is a process parameter label, which is used to describe the metallographic production process corresponding to the metallographic image data, and the text to be processed includes the metallographic process parameters.
[0016] In one embodiment of the present application, after the first model is trained based on the sample set to obtain the second model, the method further includes:
[0017] The text to be processed and the mask data to be processed are input into the second model to obtain a second metallographic image, wherein the mask data to be processed is used to mark a region of interest during the process of generating an image by the second model.
[0018] The metallographic image generating device provided in the present application comprises:
[0019] An acquisition module, used for acquiring a sample set of metallographic images, wherein the sample set includes a plurality of metallographic image data and corresponding text labels, wherein each of the plurality of metallographic image data includes a preprocessed metallographic image and mask data, and the mask data is used for marking a region of interest in the preprocessed metallographic image;
[0020] A model training module, used to train the first model based on the sample set to obtain a second model, wherein the first model is a model combining a stable diffusion model and a LoRA model, and the LoRA model is used to update the weight parameters of the cross attention layer of the stable diffusion model during the training process;
[0021] The first determination module is used to input the text to be processed into the second model to obtain a first metallographic image corresponding to the text to be processed, wherein the text to be processed includes metallographic features.
[0022] In one embodiment of the present application, in the second model, the sampling time step is positively correlated with the regional weight, and the regional weight is the weight value assigned by the second model to different regions of the first metallographic image.
[0023] In one embodiment of the present application, the pre-processed metallographic image is determined by the following steps:
[0024] Collect original metallographic images;
[0025] Clipping the original metallographic image according to a preset size to obtain a clipped original metallographic image;
[0026] The cropped original metallographic image is subjected to image enhancement processing to obtain the preprocessed metallographic image.
[0027] In one embodiment of the present application, the region of interest in the preprocessed metallographic image is: a grain image region and an inclusion image region in the preprocessed metallographic image.
[0028] In an embodiment of the present application, the mask data is binary mask data.
[0029] In one embodiment of the present application, the text label of the metallographic image data is a process parameter label, which is used to describe the metallographic production process corresponding to the metallographic image data, and the text to be processed includes the metallographic process parameters.
[0030] In one embodiment of the present application, the device further includes:
[0031] The second determination module is used to input the text to be processed and the mask data to be processed into the second model to obtain a second metallographic image, wherein the mask data to be processed is used to mark the region of interest during the process of generating an image by the second model.
[0032] The electronic device provided by the present application includes:
[0033] one or more processors;
[0034] The storage device is used to store one or more programs, and when the one or more programs are executed by the one or more processors, the electronic device implements the metallographic image generation method.
[0035] The computer-readable storage medium provided in the present application stores a computer program thereon, and when the computer program is executed by a processor of a computer, the computer is enabled to execute the metallographic image generating method.
[0036] Beneficial effects of the present application: The present application trains the first model based on the sample set to obtain the second model. The first model is a combination of the stable diffusion model and the LoRA model. The LoRA model saves training resources by fine-tuning the weight parameters of the cross-attention layer in the stable diffusion model. The present application combines the stable diffusion model and the LoRA model. In the process of generating metallographic images using the second model (i.e., the trained first model), the second model can improve the stability and detail clarity of metallographic image generation by steadily and gradually denoising; and the present application trains the first model through the above-mentioned sample set, so that the trained second model can pay more attention to the image generation of the region of interest, which is beneficial to improve the generation accuracy of the region of interest. It can be seen that the method of the present application is conducive to improving the metallographic image generation effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and together with the specification, are used to explain the principles of the present application. Obviously, the drawings described below are only some embodiments of the present application, and for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative work. In the drawings:
[0038] Figure 1 is one of the flow charts of a method for generating a metallographic image shown in an exemplary embodiment of the present application;
[0039] Figure 2 This is the second flow chart of a method for generating a metallographic image shown in an exemplary embodiment of the present application;
[0040] Figure 3 is a block diagram of a metallographic image generating device shown in an exemplary embodiment of the present application;
[0041] Figure 4 A schematic diagram of the structure of a computer system suitable for implementing an electronic device of an embodiment of the present application is shown. DETAILED DESCRIPTION
[0042] The following will describe the implementation methods of the present application with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present application from the contents disclosed in this specification. The present application can also be implemented or applied through other different specific implementation methods, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present application. It should be understood that the preferred embodiments are only for illustrating the present application, not for limiting the scope of protection of the present application.
[0043] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present application, and thus the drawings only show components related to the present application rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed at will, and the component layout may also be more complicated.
[0044] In the following description, a large number of details are discussed to provide a more thorough explanation of the embodiments of the present application. However, it is obvious to those skilled in the art that the embodiments of the present application can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present application difficult to understand.
[0045] See also Figure 1 , Figure 1 FIG. 1 is a flow chart of a method for generating a metallographic image according to an exemplary embodiment of the present application. Figure 1 As shown, in an exemplary embodiment, the metallographic image generation method includes steps S110 to S130, which are described in detail as follows:
[0046] Step S110, obtaining a sample set of metallographic images, wherein the sample set includes a plurality of metallographic image data and corresponding text labels, wherein each of the plurality of metallographic image data includes a preprocessed metallographic image and mask data, and the mask data is used to mark a region of interest in the preprocessed metallographic image;
[0047] Step S120, training the first model based on the sample set to obtain a second model, wherein the first model is a model combining a stable diffusion model and a LoRA model, and the LoRA model is used to update the weight parameters of the cross attention layer of the stable diffusion model during the training process;
[0048] Step S130: input the text to be processed into the second model to obtain a first metallographic image corresponding to the text to be processed, wherein the text to be processed includes metallographic features.
[0049] The first model is a combination of the stable diffusion model and the LoRA model. The first model is described in detail below.
[0050] The above stable diffusion model is a pre-trained Stable Diffusion basic model, such as stable-diffusion-1.5 or other Stable Diffusion series models. The pre-trained Stable Diffusion basic model is a text-to-image generation model based on deep learning.
[0051] The above-mentioned LoRA is a fine-tuning technology for neural network models, which freezes the core weights of the stable diffusion model, keeps the weights of the core part of the model unchanged during the fine-tuning process, and avoids updating the parameters of these layers; by introducing low-rank matrices in certain layers of the model to adjust the parameters, instead of retraining the entire model, thereby reducing the number of parameters that need to be updated during training. In the embodiment of the present application, the stable diffusion model and the LoRA model are combined, and the LoRA model is used to fine-tune the weight parameters of the cross-attention layer of the stable diffusion model during training. Specifically, it refers to freezing most of the weights of the Stable Diffusion base model and only training the Cross-Attention layer; and decomposing the weight matrix of the cross-attention layer in the stable diffusion model into two low-rank matrices to form a LoRA model, and only training the above-mentioned LoRA model during the first model training process. The code example is as follows:
[0052] freeze_layers=["U-Net","VAE"]
[0053] train_layers=["Cross-Attention"]
[0054] In some embodiments, in the Stable Diffusion model, the Cross-Attention layer is mainly located in the Transformer module in its U-Net architecture, so the LoRA layer can be inserted into each Transformer module, and the inserted LoRA layer is defined as two low-rank matrices obtained by decomposing the weight matrix of Cross-Attention, and all inserted LoRA layers are the above-mentioned LoRA model.
[0055] This application trains the first model based on the sample set to obtain the second model, mainly to adjust the low-rank matrix and hyperparameters. Among them, adjusting the hyperparameters means optimizing the performance of the model by selecting appropriate hyperparameters (such as learning rate, batch size, number of iterations, etc.) before model training; the weight matrix update mainly updates the low-rank matrix to update the weight matrix of the Cross-Attention layer, and the formula is as follows:
[0056] W′=W+ΔW;
[0057] ΔW=A×B
[0058] W′: weight matrix after fine-tuning; W: fixed weight matrix of the pre-trained model; A, B: low-rank matrix, obtained by decomposing the weight matrix of the Cross-Attention layer.
[0059] The above sample set includes multiple metallographic image data and corresponding text labels, each of which includes a preprocessed metallographic image and mask data. The preprocessed metallographic image is obtained by preprocessing a real metallographic image, and the mask data is used to identify the region of interest in the preprocessed metallographic image. The following is an example of the code for obtaining metallographic image data:
[0060] image=load_image("images / image1.png")
[0061] mask=load_mask("masks / image1_mask.png")
[0062] attention_input=torch.cat([image,mask],dim=0)
[0063] output=cross_attention(attention_input)
[0064] This application additionally introduces mask data to train the Cross-Attention layer of the first model, which is helpful to instruct the first model to focus on the region of interest of the input image, so that the trained second model can optimize the generation quality of the region of interest. The following is a specific description:
[0065] The calculation formula of Cross-Attention is:
[0066]
[0067] Among them, Q, K and V are query, key and value matrices respectively. The query comes from the text description, denoted by Q; the key and value come from the image features, denoted by K and V respectively. Softmax generates the attention distribution.
[0068] After introducing the mask data, the calculation formula of Cross-Attention is:
[0069]
[0070] W region(X,Y) is the region weight, determined based on the mask data, where the weight of the region of interest is higher than that of the region of no interest.
[0071] This application optimizes the generation quality of specific spatial locations (pixels or regions) based on mask data. The regional attention mechanism (with spatial attention as the core) dynamically controls the weight distribution during the generation process by adjusting the Cross-Attention layer of the model to improve the generation quality of the region of interest.
[0072] This application combines the stable diffusion model and the LoRA model. In the process of generating metallographic images using the second model (i.e., the trained first model), the second model can improve the stability and detail clarity of metallographic image generation by steadily and gradually denoising; and this application trains the first model through the above sample set, so that the trained second model can pay more attention to the image generation of the region of interest, thereby helping to improve the generation accuracy of the region of interest. It can be seen that the method of this application is conducive to improving the metallographic image generation effect.
[0073] In one embodiment of the present application, in the second model, the sampling time step is positively correlated with the regional weight, and the regional weight is the weight value assigned by the second model to different regions of the first metallographic image.
[0074] In Stable Diffusion, the sampling time step refers to the number of iterations of the denoising step in the process of generating an image, which determines the number of steps required to gradually restore a clear image from random noise. The choice of sampling time step has a direct impact on the quality and speed of generating images. More sampling steps usually generate higher quality images because the model has more opportunities to gradually remove noise.
[0075] In the embodiment of the present application, the Cross-Attention layer in the second model can determine the regional weights, and the regions with high weights are the regions of interest (key regions). By setting the sampling time step to be positively correlated with the regional weights, the time step corresponding to the regions with high weights (regions of interest) can be longer, and the time step corresponding to the regions with low weights can be shorter. This is conducive to achieving clearer details in the key regions and faster generation of other regions.
[0076] In some embodiments, the time step is determined by the following formula:
[0077] Δt t =Δt base ×(1+(α×W (x , y) ×f(t)
[0078] Among them, Δt t Indicates the time step adjusted based on the regional weight in this application; Δt baserepresents the basic time step; α represents the step adjustment amplitude; W(x,y) represents the regional weight; f(t) represents the time step change function, which can be customized.
[0079] In some embodiments, the pre-processed metallographic image is determined by the following steps:
[0080] Collect original metallographic images;
[0081] Clipping the original metallographic image according to a preset size to obtain a clipped original metallographic image;
[0082] The cropped original metallographic image is subjected to image enhancement processing to obtain the preprocessed metallographic image.
[0083] In this implementation, the preprocessed metallographic image is obtained by cropping and image enhancement processing on the original metallographic image, which is beneficial to provide high-quality input data for subsequent model training and improve the generalization ability and training efficiency of the model.
[0084] In some embodiments, the region of interest in the preprocessed metallographic image is: a grain image region and an inclusion image region in the preprocessed metallographic image.
[0085] Through the above settings, the second model trained based on the above sample set can pay more attention to the detail generation of the grain image area and the inclusion image area, so that the details of the grain image area and the inclusion image area are clearer, and other areas can be generated quickly.
[0086] In some embodiments, the mask data is binary mask data.
[0087] In this embodiment, the binary mask data is a two-dimensional array of the same size as the preprocessed metallographic image, wherein the pixel values are usually 0 or 1 (or False and True). The area with a value of 1 in the mask represents the target area of interest, while the area with a value of 0 represents the background or other unconcerned parts.
[0088] The embodiment of the present application sets the mask data as binary mask data, which helps to reduce the complexity of the mask data and helps the model pay more attention to the area of interest during the training process.
[0089] In some embodiments, the text label of the metallographic image data is a process parameter label, which is used to describe the metallographic production process corresponding to the metallographic image data, and the text to be processed includes the metallographic process parameters.
[0090] In this embodiment, the text label of the metallographic image data is a process parameter label, which describes the metallographic production process, such as production temperature, pressure, cooling rate, etc. The first model is trained by the above text labels and metallographic image data to obtain the second model. When the second model is applied, only the process parameters need to be input, and the second model can simulate the metallographic inclusion distribution and microstructure under different production conditions and output the corresponding metallographic image.
[0091] In some embodiments, after training the first model based on the sample set to obtain the second model, the method further includes:
[0092] Inputting the text to be processed and the mask data to be processed into the second model to obtain a second metallographic image;
[0093] The mask data to be processed is used to mark a region of interest during the process of generating an image by the second model.
[0094] In this implementation, the region of interest can be determined directly based on the mask data to be processed. Specifically, ControlNet can be used to load the mask data to be processed. ControlNet adds a conditional control branch to the original Stable Diffusion U-Net, processes the mask data to be processed through an additional neural network, and integrates its features into the middle layer of the main network, which enables the second model to make a clear response to the mask data to be processed during the generation process, strengthen the image detail generation of the region of interest, and quickly generate images of non-interested regions.
[0095] In the embodiment of the present application, additional mask data to be processed may be input, and the second model directly improves the image generation quality of the region of interest based on the additional input mask data to be processed.
[0096] In order to more clearly understand the technical solution of the embodiment of the present application, the following Figure 2 An exemplary description is given.
[0097] a. Prepare the sample set, including image collection and preprocessing. In the data preparation stage, for the metallographic image generation task, it is necessary to collect an appropriate amount of metallographic images; preprocess the collected metallographic images, see Figure 2 , including image cropping and data enhancement; generating regional mask data, and converting key areas (such as grains and inclusions) into binary masks; labeling the preprocessed metallographic images and their regional mask data to form a sample set.
[0098] b. Prepare and train the first model: load the trained stablediffusion model (e.g. stablediffusion 1.5); insert the LoRA layer in each transform module; define the LoRA layer as a low-rank matrix obtained by decomposing the weight matrix of Cross-Attention; freeze the other core weights of the stablediffusion model; fine-tune the LoRA layer on the sample set to learn the metallographic image features; monitor the training process to adjust the hyperparameters. At this stage, through the trained stablediffusion and fine-tuning with the prepared sample set, a LoRA model with strong adaptability and the ability to generate target image features is trained. This process requires monitoring the entire training process and gradually modifying the parameters to optimize the adaptability and generalization ability of the model. The role of the LoRA layer is to adjust the parameters of Cross-Attention by introducing a low-rank matrix, so as to adapt to the task of generating metallographic datasets without significantly changing the basic structure of the model, while significantly reducing the computational cost.
[0099] In this stage, the sample set includes preprocessed metallographic images and binary masks. Cross-Attention (or LoRA layer) is trained based on the sample set. In the early stage of model training, the metallographic image features can be learned globally. In the middle stage of model training, the sensitivity of the model to key areas can be improved. In the late stage of model training, the detail areas can be further weighted while the weight of the background areas can be weakened.
[0100] c. Train the LoRA layer in the first model through step b to obtain the second model. During the application process, the second model constructs an initial noisy image based on the input text, and then gradually denoises it (gradually converting the noisy image into a clear image. At each step, the model gradually removes the noise and gradually introduces more image structure information based on the image distribution learned during training). Combined with the trained LoRA layer, the noise is gradually converted into a real image.
[0101] d. In the practical application stage of the second model, multi-condition simulation is carried out. According to the input process parameters (such as temperature, pressure, cooling rate, etc.), the generation conditions of the LoRA layer and the stable diffusion model are adjusted to simulate the metallographic inclusion distribution and microstructure under different production conditions.
[0102] This application combines the LoRA fine-tuning technology with the stable diffusion model to accurately optimize the details in the metallographic image, effectively improve the image quality, and especially make the details more realistic. At the same time, the stable step-by-step denoising generation process ensures the stability and detail clarity of the metallographic image, avoiding the blur or distortion problems that may occur in traditional methods. Moreover, by training the Cross-Attention layer with mask data, the detail generation of the area of interest can be enhanced. In addition, LoRA fine-tuning reduces the dependence on large-scale data sets, improves training efficiency, and significantly reduces computing resource consumption.
[0103] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0104] Figure 3 FIG. 1 is a block diagram of a metallographic image generating device shown in an exemplary embodiment of the present application. Figure 3 As shown, the exemplary metallographic image generating device includes:
[0105] An acquisition module, used for acquiring a sample set of metallographic images, wherein the sample set includes a plurality of metallographic image data and corresponding text labels, wherein each of the plurality of metallographic image data includes a preprocessed metallographic image and mask data, and the mask data is used for marking a region of interest in the preprocessed metallographic image;
[0106] A model training module, used to train the first model based on the sample set to obtain a second model, wherein the first model is a model combining a stable diffusion model and a LoRA model, and the LoRA model is used to update the weight parameters of the cross attention layer of the stable diffusion model during the training process;
[0107] The first determination module is used to input the text to be processed into the second model to obtain a first metallographic image corresponding to the text to be processed, wherein the text to be processed includes metallographic features.
[0108] In one embodiment of the present application, in the second model, the sampling time step is positively correlated with the regional weight, and the regional weight is the weight value assigned by the second model to different regions of the first metallographic image.
[0109] In one embodiment of the present application, the pre-processed metallographic image is determined by the following steps:
[0110] Collect original metallographic images;
[0111] Clipping the original metallographic image according to a preset size to obtain a clipped original metallographic image;
[0112] The cropped original metallographic image is subjected to image enhancement processing to obtain the preprocessed metallographic image.
[0113] In one embodiment of the present application, the region of interest in the preprocessed metallographic image is: a grain image region and an inclusion image region in the preprocessed metallographic image.
[0114] In an embodiment of the present application, the mask data is binary mask data.
[0115] In one embodiment of the present application, the text label of the metallographic image data is a process parameter label, which is used to describe the metallographic production process corresponding to the metallographic image data, and the text to be processed includes the metallographic process parameters.
[0116] In one embodiment of the present application, the device further includes:
[0117] The second determination module is used to input the text to be processed and the mask data to be processed into the second model to obtain a second metallographic image, wherein the mask data to be processed is used to mark the region of interest during the process of generating an image by the second model.
[0118] It should be noted that the metallographic image generation device provided in the above embodiment and the metallographic image generation method provided in the above embodiment belong to the same concept, wherein the specific manner in which each module and unit performs the operation has been described in detail in the method embodiment, and will not be repeated here. In practical applications, the metallographic image generation device provided in the above embodiment can allocate the above functions to different functional modules as needed, that is, divide the internal structure of the device into different functional modules to complete all or part of the functions described above, and this is not limited here.
[0119] An embodiment of the present application also provides an electronic device, comprising: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the electronic device implements the metallographic image generation method provided in the above-mentioned embodiments.
[0120] Figure 4 The structure diagram of the computer system suitable for implementing the electronic device of the embodiment of the present application is shown. It should be noted that: Figure 4 The computer system 400 of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0121] like Figure 4As shown, the computer system 400 includes a central processing unit (CPU) 401, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 402 or the program loaded from the storage part 408 to the random access memory (RAM) 403, such as executing the method described in the above embodiment. In the RAM 403, various programs and data required for system operation are also stored. The CPU 401, ROM 402 and RAM 403 are connected to each other through a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0122] The following components are connected to the I / O interface 405: an input section 406 including a keyboard, a mouse, etc.; an output section 407 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the I / O interface 405 as needed. A removable medium 411, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 410 as needed so that a computer program read therefrom is installed into the storage section 408 as needed.
[0123] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication section 409, and / or installed from a removable medium 411. When the computer program is executed by a central processing unit (CPU) 401, various functions defined in the system of the present application are executed.
[0124] It should be noted that the computer-readable medium shown in the embodiment of the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, wherein a computer-readable computer program is carried. This propagated data signal can take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. A computer program contained on a computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0125] The flowchart and block diagram in the accompanying drawings illustrate the possible architecture, functions and operations of the system, method and computer program product according to various embodiments of the present application. Wherein, each box in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and the above-mentioned module, program segment, or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0126] The units involved in the embodiments described in this application may be implemented by software or hardware, and the units described may also be set in a processor. The names of these units do not, in some cases, constitute limitations on the units themselves.
[0127] Another aspect of the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor of a computer, causes the computer to execute the metallographic image generation method as described above. The computer-readable storage medium may be included in the electronic device described in the above embodiment, or may exist independently without being assembled into the electronic device.
[0128] Another aspect of the present application also provides a computer program product or a computer program, which includes a computer instruction stored in a computer-readable storage medium. A processor of a computer device reads the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device executes the metallographic image generation method provided in each of the above embodiments.
[0129] The above embodiments are merely illustrative of the principles and effects of the present application, and are not intended to limit the present application. Anyone familiar with the technology may modify or change the above embodiments without violating the spirit and scope of the present application. Therefore, all equivalent modifications or changes made by a person of ordinary skill in the art without departing from the spirit and technical ideas disclosed in the present application shall still be covered by the claims of the present application.
Claims
1. A metallographic image generation method, characterized in that: include: Acquire a sample set of metallographic images, wherein the sample set includes a plurality of metallographic image data and corresponding text labels, wherein each of the plurality of metallographic image data includes a preprocessed metallographic image and mask data, and the mask data is used to mark a region of interest in the preprocessed metallographic image; Training the first model based on the sample set to obtain a second model, wherein the first model is a model combining a stable diffusion model and a LoRA model, and the LoRA model is used to update the weight parameters of the cross attention layer of the stable diffusion model during the training process; The text to be processed is input into the second model to obtain a first metallographic image corresponding to the text to be processed, wherein the text to be processed includes metallographic features.
2. The method according to claim 1, characterized in that In the second model, the sampling time step is positively correlated with the regional weight, and the regional weight is the weight value assigned by the second model to different regions of the first metallographic image.
3. The method according to claim 1, characterized in that The pre-processed metallographic image is determined by the following steps: Collect original metallographic images; Clipping the original metallographic image according to a preset size to obtain a clipped original metallographic image; The cropped original metallographic image is subjected to image enhancement processing to obtain the preprocessed metallographic image.
4. The method according to claim 1, characterized in that The regions of interest in the preprocessed metallographic image are: a grain image region and an inclusion image region in the preprocessed metallographic image.
5. The method according to claim 1, characterized in that The mask data is binary mask data.
6. The method according to claim 1, characterized in that The text label of the metallographic image data is a process parameter label, which is used to describe the metallographic production process corresponding to the metallographic image data, and the text to be processed includes the metallographic process parameters.
7. The method according to claim 1, characterized in that After the first model is trained based on the sample set to obtain the second model, the method further includes: Inputting the text to be processed and the mask data to be processed into the second model to obtain a second metallographic image; The mask data to be processed is used to mark a region of interest during the process of generating an image by the second model.
8. A metallographic image generating device, characterized in that: include: An acquisition module, used for acquiring a sample set of metallographic images, wherein the sample set includes a plurality of metallographic image data and corresponding text labels, wherein each of the plurality of metallographic image data includes a preprocessed metallographic image and mask data, and the mask data is used for marking a region of interest in the preprocessed metallographic image; A model training module, used to train the first model based on the sample set to obtain a second model, wherein the first model is a model combining a stable diffusion model and a LoRA model, and the LoRA model is used to update the weight parameters of the cross attention layer of the stable diffusion model during the training process; The first determination module is used to input the text to be processed into the second model to obtain a first metallographic image corresponding to the text to be processed, wherein the text to be processed includes metallographic features.
9. A device, characterized in that: include: one or more processors and memory, A computer program is stored in the memory, and when the one or more processors execute the computer program, the device executes the metallographic image generating method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, which, when executed by one or more processors, enables the device to execute the metallographic image generating method according to any one of claims 1 to 7.
Citation Information
Cited By
NDVI product generation method and device based on diffusion model, equipment and medium
CN121213392A