A method for generating crystal image samples based on a diffusion model

By using a diffusion model-based crystal image generation method, the problems of high cost and long experimental cycle in crystal image generation are solved, and high-quality crystal growth images can be generated quickly, meeting the research needs of researchers for the diversity of crystal morphology.

CN120635234BActive Publication Date: 2026-05-26融域智慧(西安)智能科技有限公司
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
融域智慧(西安)智能科技有限公司
Filing Date
2025-05-28
Publication Date
2026-05-26

Smart Images

  • Figure CN120635234B_ABST
    Figure CN120635234B_ABST
Patent Text Reader

Abstract

This invention discloses a method for generating crystal image samples based on a diffusion model, comprising: generating a standard comparison image based on a reaction flask photograph; obtaining crystal identification features based on the shape and color of the crystal; detecting the presence of the crystal by using the crystal identification features and the standard comparison image; and generating a crystal image based on the crystal presence result. The crystal image generation is performed using a diffusion model. This invention, through an AI method based on a diffusion model, can generate crystal growth images under different conditions within seconds, without requiring long waiting times for experimental results or physical experiments. It can generate high-quality crystal growth samples, suitable for materials research and educational visualization. The shape, color, transparency, and other parameters of the crystal can be adjusted to match experimental data of specific chemical systems. It utilizes AI to generate large-scale, high-quality crystal growth image datasets, assisting the application of deep learning in materials science.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image generation technology, and in particular to a method for generating crystal image samples based on a diffusion model. Background Technology

[0002] In the fields of materials science and chemistry, crystal images serve as crucial evidence for studying crystal structure, growth mechanisms, and properties. They provide a direct visual representation of crystal morphology under different conditions, offering vital data support for new material development and the exploration of chemical reaction mechanisms. By analyzing crystal images, researchers can gain a deeper understanding of crystal microstructure, thereby optimizing material properties and chemical reaction processes.

[0003] Currently, acquiring crystal images primarily relies on real chemical experiments. This involves setting specific experimental conditions to observe and record the crystal growth process, and then using equipment such as microscopes and scanning electron microscopes to acquire crystal images. This experiment-based image acquisition method is traditional and reliable, and has accumulated a large amount of crystal image data in past research.

[0004] However, this technology has significant limitations. On the one hand, real chemical experiments require sophisticated equipment, reagents, and highly skilled personnel, leading to high costs for crystal image generation. Simultaneously, crystal growth processes often take a considerable amount of time, resulting in long experimental cycles and significantly limiting research efficiency. On the other hand, the growth morphology of crystals under different temperatures, pressures, and concentrations exhibits high complexity and variability. Relying solely on chemical experiments makes it difficult to simulate crystal growth under various conditions in a short time, failing to meet researchers' urgent need for diverse crystal morphologies and becoming a technological bottleneck restricting the rapid development of materials science and chemistry. Summary of the Invention

[0005] This invention provides a method for generating crystal image samples based on a diffusion model. This method addresses the high costs associated with generating crystal images in existing technologies, where real chemical experiments require advanced equipment, reagents, and skilled personnel, leading to long crystal growth processes and extended experimental cycles that significantly limit research efficiency. Furthermore, the growth morphology of crystals under different temperatures, pressures, and concentrations exhibits high complexity and variability. Relying solely on chemical experiments makes it difficult to simulate crystal growth under various conditions in a short time, thus failing to meet researchers' urgent need for studying the diversity of crystal morphologies.

[0006] On one hand, embodiments of the present invention provide a method for generating crystal image samples based on a diffusion model, comprising:

[0007] Generate a standard comparison chart based on the reaction flask photograph;

[0008] Crystal identification features are derived from the shape and color of the crystal;

[0009] The presence of a crystal is obtained by detecting the presence of a crystal using the crystal identification features and the standard comparison diagram.

[0010] A crystal image is generated based on the presence of the crystal.

[0011] The crystal image is generated using a diffusion model.

[0012] In one possible implementation, generating a standard comparison image based on the reaction flask photograph includes:

[0013] The reaction flask photograph was preprocessed to obtain the standard comparison image without background.

[0014] In one possible implementation, the step of obtaining the crystal presence result by detecting the crystal presence through the crystal identification features and the standard comparison map includes:

[0015] The presence of the crystal is obtained by using the standard comparison chart and the crystal recognition features as sample recognition images;

[0016] The presence of crystals includes both the presence of crystals and the absence of crystals.

[0017] In one possible implementation, generating the crystal image based on the presence of the crystal includes:

[0018] A crystal image is generated based on the crystal type from the reagent bottle image where the crystal presence result is indicated.

[0019] A crystal image is obtained from the reagent bottle image where the presence of crystals results in the absence of crystals by using shape and color constraints.

[0020] In one possible implementation, the crystal image generation based on the crystal type in the reagent bottle image where the crystal presence result is the presence of crystals is achieved by filling in the presence of crystals using a diffusion model and inpainting techniques.

[0021] In one possible implementation, the crystal image is obtained from the crystalless reagent bottle image by means of shape and color constraints through a diffusion model and ControlNet technology.

[0022] In one possible implementation, the crystal image generation is performed via a diffusion model and further includes;

[0023] The LoRA model is used to constrain the shape and color conditions for generating crystal images.

[0024] In one possible implementation, the crystal image generation via a diffusion model further includes:

[0025] The crystal image is then rendered in HDR.

[0026] The crystal image sample generation method based on a diffusion model in this invention has the following advantages:

[0027] (1) By using an AI method based on a diffusion model, crystal growth images under different conditions can be generated in seconds without having to wait for experimental results for a long time.

[0028] (2) High-quality crystal growth samples can be generated without physical experiments, which is suitable for materials research and educational visualization.

[0029] (3) The shape, color, transparency and other parameters of the crystal can be adjusted to match the experimental data of a specific chemical system.

[0030] (4) Use AI to generate large-scale, high-quality crystal growth image datasets to assist the application of deep learning in materials science. Attached Figure Description

[0031] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0032] Figure 1 A flowchart illustrating a method for generating crystal image samples based on a diffusion model, provided in this application embodiment;

[0033] Figure 2 This is a flowchart illustrating a method for generating crystal image samples based on a diffusion model, as provided in an embodiment of this application. Detailed Implementation

[0034] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0035] Figure 1A flowchart illustrating a crystal image sample generation method based on a diffusion model, provided in an embodiment of the present invention; the embodiment of the present invention provides a crystal image sample generation method based on a diffusion model, comprising:

[0036] Generate a standard comparison chart based on the reaction flask photograph;

[0037] Crystal identification features are derived from the shape and color of the crystal;

[0038] The presence of a crystal is obtained by detecting the presence of a crystal using the crystal identification features and the standard comparison diagram.

[0039] A crystal image is generated based on the presence of the crystal.

[0040] The crystal image is generated using a diffusion model.

[0041] The generation of a standard comparison image based on the reaction flask photograph includes:

[0042] The reaction flask photograph was preprocessed to obtain the standard comparison image without background.

[0043] The crystal presence detection method, which uses the crystal identification features and the standard comparison image to obtain the crystal presence result, includes:

[0044] The presence of the crystal is obtained by using the standard comparison chart and the crystal recognition features as sample recognition images;

[0045] The presence of crystals includes both the presence of crystals and the absence of crystals.

[0046] The step of generating a crystal image based on the presence of the crystal includes:

[0047] A crystal image is generated based on the crystal type from the reagent bottle image where the crystal presence result is indicated.

[0048] A crystal image is obtained from the reagent bottle image where the presence of crystals results in the absence of crystals by using shape and color constraints.

[0049] The crystal image is generated based on the crystal type in the reagent bottle image where the crystal is present, by filling in the present crystal using a diffusion model and inpainting technology.

[0050] The crystal image, obtained by constraining the shape and color of the reagent bottle image where the presence of crystals results in the absence of crystals, is generated using a diffusion model and ControlNet technology to create the absence of crystals.

[0051] The crystal image generation, performed via a diffusion model, also includes;

[0052] The LoRA model is used to constrain the shape and color conditions for generating crystal images.

[0053] The crystal image generation, performed via a diffusion model, also includes:

[0054] The crystal image is then rendered in HDR.

[0055] For example, such as Figure 1 and 2 As shown, the photo of the reagent bottle for the reaction is first input as the image of the crystal to be generated. Then, the background of the bottle is cleaned through preprocessing. Then, the presence of crystal is identified by the shape and color parameters of the desired crystal. If the crystal is detected in the target image, the scheme of generating a new crystal on the existing crystal is executed. If the crystal is not detected in the target image, the scheme of directly generating the crystal is executed.

[0056] The method for generating new crystals on existing crystals utilizes a diffusion model combined with inpainting technology to control the gradual filling of the existing crystal and control the growth size. Simultaneously, LoRA is used to handle material and color. Pre-trained LoRA models for different types of crystal colors and textures, such as "blue copper sulfate crystals," are used to select the appropriate LoRA model for different shapes and colors.

[0057] Direct crystal generation involves processing the shape using a diffusion model and ControlNet based on input parameters: the input is a crystal morphology diagram obtained from physical simulation, ensuring that the AI-generated crystal shape matches the experimental results. Simultaneously, LoRA is used to process the material and color.

[0058] In one possible implementation, the core implementation logic of crystal presence detection, diffusion model combined with Inpainting / ControlNet, LoRA color control, and HDR rendering is illustrated through code examples as follows:

[0059] importtorch

[0060] importtorch.nnasnn

[0061] from ultralyticsimportYOLO

[0062] fromtorchvisionimporttransforms

[0063] fromPILimportImage

[0064] import numpyasnp

[0065] #---------------------------

[0066] #1. Data Preprocessing and Feature Fusion

[0067] #---------------------------

[0068] classCrystalFeatureProcessor:

[0069] def__init__(self):

[0070] self.transform=transforms.Compose([

[0071] transforms.Resize((640,640)), #YOLOv8 default input size

[0072] transforms.ToTensor(),

[0073] transforms.Normalize(mean=[0.485,0.456,0.406],std=[0.229,0.224,0.225]) ])

[0075] defprocess(self,img_path,shape_feat,color_feat):

[0076] """

[0077] Input: Path to the reaction flask image, shape features (dictionary), color features (dictionary)

[0078] Output: The fused model input (image tensor + feature vector)

[0079] """

[0080] #Preprocessed image (after background removal)

[0081] img=Image.open(img_path).convert('RGB')

[0082] img_tensor=self.transform(img).unsqueeze(0)#[1,3,640,640]

[0083] #Encode shape / color features as a 12-dimensional vector (example dimension)

[0084] feat_vec=torch.tensor([

[0085] shape_feat['aspect_ratio'], # aspect ratio

[0086] shape_feat['roundness'],#roundness

[0087] shape_feat['convex_area'], # Ratio of convex hull area

[0088] color_feat['hue_mean'], # Hue mean

[0089] color_feat['sat_std'], # Saturation variance

[0090] color_feat['val_mean']# Average brightness

[0091] ]).float().unsqueeze(0)#[1,6] (This example is simplified to 6 dimensions, but can be expanded in practice)

[0092] returnimg_tensor,feat_vec

[0093] #---------------------------

[0094] #2. Improved YOLOv8 object detection model

[0095] #---------------------------

[0096] classCrystalDetector(nn.Module):

[0097] def__init__(self,pretrained=True):

[0098] super().__init__()

[0099] #Load the YOLOv8n base model (lightweight, suitable for small goals)

[0100] self.yolo=YOLO('yolov8n.pt')ifpretrainedelseYOLO('yolov8n.yaml')

[0101] #Add a new feature fusion layer (maps 6-dimensional features to image feature dimensions)

[0102] self.feat_fusion=nn.Sequential(

[0103] nn.Linear(6,256), # Maps to the same number of channels (256) as the YOLO feature map

[0104] nn.ReLU() )

[0106] defforward(self,img_tensor,feat_vec):

[0107] #YOLO basic feature extraction (output feature map: [1,256,80,80])

[0108] yolo_feats=self.yolo.model(img_tensor)[0]#Get backbone network output

[0109] #Fusing Feature Vectors: Expanding a 12-dimensional vector to the same size as the feature map

[0110] batch,channels,height,width=yolo_feats.shape

[0111] fused_feats=self.feat_fusion(feat_vec).reshape(batch,-1,1,1)#[1,256,1,1]

[0112] fused_feats=fused_feats.repeat(1,1,height,width)#[1,256,80,80]

[0113] #Feature maps are added and merged element by element

[0114] final_feats=yolo_feats+fused_feats

[0115] # Output results via YOLO detector head (bounding box + class probability)

[0116] detections=self.yolo.head(final_feats)

[0117] Returndetections

[0118] #---------------------------

[0119] #3. Model Training Example (Simplified Process)

[0120] #---------------------------

[0121] if__name__=="__main__":

[0122] #Initialize processor and model

[0123] processor=CrystalFeatureProcessor()

[0124] detector=CrystalDetector(pretrained=True)

[0125] #Simulated Input (Example Data)

[0126] img_tensor,feat_vec=processor.process(

[0127] 'example_bottle.jpg',

[0128] shape_feat={'aspect_ratio':1.2,'roundness':0.8,'convex_area':0.95},

[0129] color_feat={'hue_mean':200,'sat_std':15,'val_mean':180} )

[0131] #Forward Reasoning

[0132] detections=detector(img_tensor,feat_vec)

[0133] print(f"Detection result shape: {detections.shape}") # Output: [1,8400,85] (YOLOv8 standard output format)

[0134] importtorch

[0135] fromdiffusersimportStableDiffusionControlNetPipeline,ControlNetModel

[0136] fromPILimportImage

[0137] #---------------------------

[0138] #1. Initialize ControlNet pipeline (morphological graph constraints)

[0139] #---------------------------

[0140] classCrystalControlNetGenerator:

[0141] def__init__(self,controlnet_path='controlnet_crystal_shape.pt'):

[0142] #Load the ControlNet model (pre-trained crystal morphology control network)

[0143] self.controlnet=ControlNetModel.from_pretrained(

[0144] controlnet_path,

[0145] torch_dtype=torch.float16

[0146] .to("cuda")

[0147] #Load the StableDiffusion main model

[0148] self.pipeline=StableDiffusionControlNetPipeline.from_pretrained(

[0149] "runwayml / stable-diffusion-v1-5",

[0150] controlnet = self.controlnet,

[0151] torch_dtype=torch.float16

[0152] .to("cuda")

[0153] defgenerate(self,morph_map,color_feat,shape_type='cube'):

[0154] """

[0155] Input: Physical simulation morphology map (edge ​​map), color features, crystal shape type

[0156] Output: Generated crystal image

[0157] """

[0158] #Preprocess the morphology map (convert it to an edge map)

[0159] morph_map = morph_map.convert('L') # Grayscale image

[0160] morph_map=morph_map.resize((512,512))

[0161] #Build prompt words

[0162] prompt=f"realistic{color_feat['color_name']}{shape_type}crystal,matchingthegivenmorphology"

[0163] #Generate controls (shape constraints via ControlNet)

[0164] output = self.pipeline(

[0165] prompt=prompt,

[0166] image=morph_map, # Morphological map used as control signal

[0167] num_inference_steps=50,

[0168] guidance_scale=7.5

[0169] ).images[0]

[0170] returnoutput

[0171] #---------------------------

[0172] #2. Usage Example

[0173] #---------------------------

[0174] if__name__=="__main__":

[0175] generator=CrystalControlNetGenerator(controlnet_path='controlnet_crystal_shape.pt')

[0176] #Load the physical simulation morphology map (edge ​​map generated by molecular dynamics simulation)

[0177] morph_map=Image.open('morphology_map.png').convert('RGB')

[0178] #Color characteristics (white sodium chloride)

[0179] color_feat={'color_name':'white','hue':0}

[0180] #Generate cubic crystals

[0181] output_img=generator.generate(morph_map,color_feat,shape_type='cube')

[0182] output_img.save('generated_crystal_controlnet.jpg')

[0183] importtorch

[0184] fromdiffusersimportStableDiffusionInpaintPipeline

[0185] fromtorchvisionimporttransforms

[0186] #---------------------------

[0187] #1. Initialize the Inpainting pipeline (with LoRA color control)

[0188] #---------------------------

[0189] classCrystalInpainter:

[0190] def__init__(self,lora_path='lora_cu_so4.pt'):

[0191] #Load the StableDiffusionInpainting pre-trained model

[0192] self.pipeline=StableDiffusionInpaintPipeline.from_pretrained(

[0193] "runwayml / stable-diffusion-inpainting",

[0194] torch_dtype=torch.float16

[0195] .to("cuda")

[0196] #Load LoRA adapter (color / material control)

[0197] self.pipeline.unet.load_attn_procs(lora_path) # Load the pre-trained copper sulfate LoRA model

[0198] defgenerate(self,input_img,mask,shape_feat,color_feat,size_scale=1.5):

[0199] """

[0200] Input: Original image (including existing crystal), mask (marking the area to be filled), shape features, color features, growth size scale

[0201] Output: The generated crystal image

[0202] """

[0203] #Preprocessing Input

[0204] input_img=transforms.ToTensor()(input_img).unsqueeze(0).to("cuda")#[1,3,H,W]

[0205] mask=transforms.ToTensor()(mask).unsqueeze(0).to("cuda")#[1,1,H,W]

[0206] #Constructing cue words (combining shape and color features)

[0207] prompt=f"realistic{color_feat['color_name']}crystal,{shape_feat['shape_type']},sizescale{size_scale}"

[0208] #Generation control: Limit noise to be added only to the mask area

[0209] self.pipeline.scheduler.set_timesteps(1000) # Increase the number of diffusion steps to improve detail

[0210] withtorch.no_grad():

[0211] output = self.pipeline(

[0212] prompt=prompt,

[0213] image=input_img,

[0214] mask_image=mask,

[0215] guidance_scale=7.5,

[0216] num_inference_steps=1000

[0217] ).images[0]

[0218] returnoutput

[0219] #---------------------------

[0220] #2. Usage Example

[0221] #---------------------------

[0222] if__name__=="__main__":

[0223] inpainter=CrystalInpainter(lora_path='lora_cu_so4.pt')

[0224] #Load input image and mask (existing crystal regions are preserved, and edge regions are marked as to be filled)

[0225] input_img=Image.open('existing_crystal.jpg').convert('RGB')

[0226] mask = Image.open('inpainting_mask.png').convert('L') # The white area is to be filled.

[0227] #Assuming shape characteristics (needle-like) and color characteristics (blue copper sulfate)

[0228] shape_feat={'shape_type':'acicular'}#needle-shaped

[0229] color_feat={'color_name':'blue','hue':200}

[0230] #Generate a new crystal (size magnified 1.5 times)

[0231] output_img=inpainter.generate(input_img,mask,shape_feat,color_feat,size_scale=1.5)

[0232] output_img.save('generated_crystal_inpainting.jpg')

[0233] importcv2

[0234] import numpyasnp

[0235] defhdr_render(img_path,output_path):

[0236] """

[0237] Input: Path to the generated crystal image

[0238] Output: Image rendered in HDR

[0239] """

[0240] # Read the image and extend the dynamic range (0-255→0-4096)

[0241] img=cv2.imread(img_path).astype(np.float32)

[0242] hdr_img=img (4096 / 255)# Expanded to 12-bit HDR

[0243] #Tone Mapping (Reinhard Algorithm)

[0244] tonemap=cv2.createTonemapReinhard(gamma=1.5,intensity=0.5)

[0245] ldr_img=tonemap.process(hdr_img)

[0246] #Material Enhancement (Adding Specular Reflection)

[0247] gray=cv2.cvtColor(ldr_img,cv2.COLOR_BGR2GRAY)

[0248] _,mask=cv2.threshold(gray,0.8,1.0,cv2.THRESH_BINARY) #The highlighted area is used as the reflection area.

[0249] ldr_img += mask[...,None] 0.1# Increase reflectivity

[0250] #Convert back to 8 bits and save

[0251] ldr_img=np.clip(ldr_img 255,0,255).astype(np.uint8)

[0252] cv2.imwrite(output_path,ldr_img)

[0253] #Usage Example

[0254] if__name__=="__main__":

[0255] hdr_render('generated_crystal.jpg','hdr_crystal.jpg')

[0256] importtorch

[0257] frompeftimportLoraConfig,get_peft_model

[0258] fromdiffusersimportStableDiffusionPipeline

[0259] #---------------------------

[0260] #1. LoRA Adapter Definition and Training

[0261] #---------------------------

[0262] deftrain_lora_crystal(

[0263] model_path='runwayml / stable-diffusion-v1-5',

[0264] train_data_path='crystal_dataset / ',

[0265] output_path='lora_crystal.pt',

[0266] epochs=10 ):

[0268] #Load the base model

[0269] pipe=StableDiffusionPipeline.from_pretrained(model_path,torch_dtype=torch.float16).to("cuda")

[0270] # Define LoRA configuration (fine-tune the cross-attention layer only)

[0271] lora_config=LoraConfig(

[0272] r=8, # Rank of a low-rank matrix

[0273] lora_alpha=16,

[0274] target_modules=["to_q","to_k"], # Only fine-tune the Q and K matrices

[0275] lora_dropout=0.05,

[0276] bias="none",

[0277] task_type="TEXT_TO_IMAGE" )

[0279] #Add LoRA adapter to UNet

[0280] pipe.unet=get_peft_model(pipe.unet,lora_config)

[0281] pipe.unet.print_trainable_parameters() # Outputs trainable parameters (approximately 1%)

[0282] #Simulated training loop (actual training requires loading the crystal dataset)

[0283] optimizer=torch.optim.AdamW(pipe.unet.parameters(),lr=1e-4)

[0284] forepochinrange(epochs):

[0285] #Assuming data_loader provides (hint words, images, color labels)

[0286] forbatchindata_loader:

[0287] prompts,images,color_labels=batch

[0288] #Forward generation

[0289] latents=pipe(prompts,return_dict=False)[0]

[0290] # Calculate the loss (L1 loss + color feature matching loss)

[0291] loss=torch.nn.functional.l1_loss(latents,images)

[0292] #Reverse propagation

[0293] loss.backward()

[0294] optimizer.step()

[0295] optimizer.zero_grad()

[0296] #Save LoRA adapter

[0297] pipe.unet.save_attn_procs(output_path)

[0298] #---------------------------

[0299] #2. Training Startup Example

[0300] #---------------------------

[0301] if__name__=="__main__":

[0302] train_lora_crystal(

[0303] model_path='runwayml / stable-diffusion-v1-5',

[0304] train_data_path='crystal_dataset / ',

[0305] output_path='lora_cu_so4.pt', # Copper sulfate crystals LoRA

[0306] epochs=10 )

[0308] In one possible embodiment, the Python implementation code for a YOLOv8-based crystal presence recognition target detection model includes data preprocessing, feature fusion, model definition, and inference examples:

[0309] importtorch

[0310] import numpyasnp

[0311] fromPILimportImage

[0312] from ultralyticsimportYOLO

[0313] fromtorchvisionimporttransforms

[0314] fromscipy.ndimageimportgaussian_filter

[0315] #---------------------------

[0316] #1. Crystal Feature Extraction Tools

[0317] #---------------------------

[0318] classCrystalFeatureExtractor:

[0319] def__init__(self):

[0320] #Color Space Conversion Tool

[0321] self.rgb2hsv=transforms.Compose([

[0322] transforms.ToTensor(),

[0323] lambdax:x.permute(1,2,0).numpy(),#HWC format

[0324] lambdax:np.array(Image.fromarray((x 255).astype(np.uint8)).convert('HSV')) / 255.0 # Normalize HSV to [0,1] ])

[0326] defextract_shape_features(self,img_mask):

[0327] """

[0328] Input: Crystal binary mask (0-1, 1 represents the crystal region)

[0329] Output: Shape feature dictionary (aspect ratio, roundness, convex hull area ratio)

[0330] """

[0331] #Calculate the bounding box

[0332] y,x=np.where(img_mask>0.5)

[0333] iflen(y)==0:#No crystals

[0334] return{'aspect_ratio':0.0,'roundness':0.0,'convex_area_ratio':0.0}

[0335] min_x, max_x = x.min(), x.max()

[0336] min_y, max_y = y.min(), y.max()

[0337] width=max_x-min_x+1

[0338] height = max_y - min_y + 1

[0339] aspect_ratio = width / height if height != 0 else 0.0 # Aspect ratio

[0340] # Calculate roundness (4π area / perimeter²)

[0341] area = np.sum(img_mask)

[0342] perimeter=self._calculate_perimeter(img_mask)

[0343] roundness=(4 np.pi area) / (perimeter 2) if perimeter != 0 else 0.0

[0344] # Calculate the convex hull area ratio (convex hull area / actual area)

[0345] convex_area=self._calculate_convex_area(img_mask)

[0346] convex_area_ratio=area / convex_areaifconvex_area!=0else0.0

[0347] return{

[0348] 'aspect_ratio':aspect_ratio,

[0349] 'roundness':roundness,

[0350] 'convex_area_ratio':convex_area_ratio

[0351] }

[0352] defextract_color_features(self,img_rgb,img_mask):

[0353] """

[0354] Input: RGB image (HWC, 0-1), crystal binary mask (0-1)

[0355] Output: Color feature dictionary (hue mean, saturation variance, lightness mean)

[0356] """

[0357] img_hsv=self.rgb2hsv(Image.fromarray((img_rgb 255).astype(np.uint8)))#HSV[0,1]

[0358] crystal_h=img_hsv[...,0][img_mask>0.5]

[0359] crystal_s=img_hsv[...,1][img_mask>0.5]

[0360] crystal_v=img_hsv[...,2][img_mask>0.5]

[0361] iflen(crystal_h)==0:#No crystal

[0362] return{'hue_mean':0.0,'sat_std':0.0,'val_mean':0.0}

[0363] return{

[0364] 'hue_mean': np.mean(crystal_h), # Hue mean (0-1)

[0365] 'sat_std':np.std(crystal_s), #saturation variance

[0366] 'val_mean': np.mean(crystal_v) # Mean value of brightness (0-1)

[0367] }

[0368] def_calculate_perimeter(self,mask):

[0369] """Calculate the perimeter of the mask (using edge detection)""

[0370] mask_edges=gaussian_filter(mask,sigma=0.5) # Smooth edges

[0371] dx=np.gradient(mask_edges,axis=1)

[0372] dy=np.gradient(mask_edges,axis=0)

[0373] perimeter = np.sum(np.sqrt(dx) 2+dy 2))

[0374] returnperimeter

[0375] def_calculate_convex_area(self,mask):

[0376] Calculating the area of ​​the convex hull (simplified implementation)

[0377] fromskimage.morphologyimportconvex_hull_image

[0378] convex_mask=convex_hull_image(mask.astype(bool))

[0379] return np.sum(convex_mask)

[0380] #---------------------------

[0381] #2. Improved YOLOv8 Detection Model (Feature Fusion Version)

[0382] #---------------------------

[0383] class CrystalDetectionModel:

[0384] def__init__(self,pretrained=True):

[0385] #Load the YOLOv8n base model (lightweight, suitable for small goals)

[0386] self.model=YOLO('yolov8n.pt'ifpretrainedelse'yolov8n.yaml')

[0387] #Feature preprocessing (mapping 6-dimensional features to YOLO feature dimensions)

[0388] self.feat_proj=torch.nn.Sequential(

[0389] torch.nn.Linear(6,256), #6-dimensional feature → 256-dimensional (matches the number of channels in the YOLO neck feature)

[0390] torch.nn.ReLU()

[0391] ).to(self.model.device)

[0392] def_fuse_features(self,yolo_feats,shape_color_feats):

[0393] """Fusing shape and color features with YOLO image features""

[0394] batch_size,channels,height,width=yolo_feats.shape

[0395] #Project the feature vectors and expand them to the feature map size

[0396] proj_feats=self.feat_proj(shape_color_feats).reshape(batch_size,channels,1,1)

[0397] proj_feats=proj_feats.repeat(1,1,height,width)#[B,256,H,W]

[0398] #Feature addition and fusion

[0399] returnyolo_feats+proj_feats

[0400] deftrain(self,train_dataset,epochs=10,lr=1e-3):

[0401] """Training the Model (Simplified Interface)"""

[0402] #Freeze the first 3 layers of the YOLO backbone network (preserve general features)

[0403] fori,layerinumerate(self.model.model[:3]):

[0404] forparaminlayer.parameters():

[0405] param.requires_grad=False

[0406] # Define the optimizer (train only the feature fusion layer and detection head)

[0407] optimizer = torch.optim.Adam(

[0408] list(self.model.model.parameters())+list(self.feat_proj.parameters()),

[0409] lr=lr )

[0411] #Training loop (actually requires a data loader)

[0412] forepochinrange(epochs):

[0413] forimgs, masks, labels intrain_dataset: # Assume data format: images, masks, labels

[0414] #Extract shape and color features

[0415] extractor=CrystalFeatureExtractor()

[0416] feats_list=[]

[0417] forimg,maskinzip(imgs,masks):

[0418] shape_feats=extractor.extract_shape_features(mask)

[0419] color_feats=extractor.extract_color_features(img,mask)

[0420] # Concatenate into a 6-dimensional feature vector (aspect, round, convex, hue, sat_std, val)

[0421] feat_vec=torch.tensor([

[0422] shape_feats['aspect_ratio'],

[0423] shape_feats['roundness'],

[0424] shape_feats['convex_area_ratio'],

[0425] color_feats['hue_mean'],

[0426] color_feats['sat_std'],

[0427] color_feats['val_mean']

[0428] ]).to(self.model.device)

[0429] feats_list.append(feat_vec)

[0430] shape_color_feats=torch.stack(feats_list)#[B,6]

[0431] #Forward propagation (obtaining YOLO backbone features)

[0432] yolo_feats=self.model.model(imgs)[0]# Backbone output feature map[B,256,H,W]

[0433] #Feature Fusion

[0434] fused_feats=self._fuse_features(yolo_feats,shape_color_feats)

[0435] #Prediction via YOLO detector head

[0436] outputs=self.model.head(fused_feats)

[0437] # Calculate the loss (using YOLO's built-in loss function)

[0438] loss=self.model.loss(outputs,labels)

[0439] #Reverse propagation

[0440] optimizer.zero_grad()

[0441] loss.backward()

[0442] optimizer.step()

[0443] print(f"Epoch{epoch+1} / {epochs},Loss:{loss.item():.4f}")

[0444] defpredict(self,img_path,threshold=0.5):

[0445] """Inference prediction (returns the existence of a crystal)"""

[0446] #Load and preprocess the image

[0447] img=Image.open(img_path).convert('RGB')

[0448] img_tensor=self.model.preprocess(img) #YOLO built-in preprocessing (resize + normalization)

[0449] #When there is no real mask, the entire image is used as the candidate region (in practice, the mask needs to be obtained first through a segmentation model).

[0450] dummy_mask=np.ones((img.height,img.width)) # Example using full-image mask

[0451] extractor=CrystalFeatureExtractor()

[0452] shape_feats=extractor.extract_shape_features(dummy_mask)

[0453] color_feats=extractor.extract_color_features(np.array(img) / 255.0,dummy_mask)

[0454] feat_vec=torch.tensor([

[0455] shape_feats['aspect_ratio'],

[0456] shape_feats['roundness'],

[0457] shape_feats['convex_area_ratio'],

[0458] color_feats['hue_mean'],

[0459] color_feats['sat_std'],

[0460] color_feats['val_mean']

[0461] ]).to(self.model.device).unsqueeze(0)#[1,6]

[0462] #Forward Reasoning

[0463] yolo_feats=self.model.model(img_tensor)[0]

[0464] fused_feats=self._fuse_features(yolo_feats,feat_vec)

[0465] outputs=self.model.head(fused_feats)

[0466] #Analysis of test results (presence of crystals)

[0467] detections=self.model.postprocess(outputs,img.shape[:2])

[0468] has_crystal=any(det['conf']>thresholdfordetindetections)

[0469] return has_crystal

[0470] #---------------------------

[0471] #3. Usage Example

[0472] #---------------------------

[0473] if__name__=="__main__":

[0474] # Initialize the model (load pre-trained weights)

[0475] detector=CrystalDetectionModel(pretrained=True)

[0476] #Test Reasoning (Photo of a reaction flask assuming the presence of crystals)

[0477] test_img_path='test_crystal_bottle.jpg'

[0478] has_crystal=detector.predict(test_img_path)

[0479] print(f"Detection result: {'Crystal present' if has_crystalelse 'No crystal present'}")

[0480] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0481] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A method for generating crystal image samples based on a diffusion model, characterized in that, include: Generate a standard comparison chart based on the reaction flask photograph; Crystal identification features are derived from the shape and color of the crystal; The presence of a crystal is obtained by detecting the presence of a crystal using the crystal identification features and the standard comparison diagram. A crystal image is generated based on the presence of the crystal. The crystal image is generated using a diffusion model; The generation of the standard comparison chart based on the reaction flask photograph includes: The reaction flask photographs were preprocessed to obtain the standard comparison image without background. The crystal presence detection method, which uses the crystal identification features and the standard comparison image to obtain the crystal presence result, includes: The presence of the crystal is obtained by using the standard comparison chart and the crystal recognition features as sample recognition images; The results of the crystal presence include the presence of crystals and the absence of crystals; The step of generating a crystal image based on the presence of the crystal includes: A crystal image is generated based on the crystal type from the reagent bottle image where the crystal presence result is indicated. A crystal image is obtained from the reagent bottle image where the presence of crystals results in the absence of crystals by using shape and color constraints.

2. The crystal image sample generation method based on a diffusion model according to claim 1, characterized in that, The crystal image is generated based on the crystal type in the reagent bottle image where the crystal is present, by filling in the present crystal using a diffusion model and inpainting technology.

3. The crystal image sample generation method based on a diffusion model according to claim 2, characterized in that, The crystal image, obtained by constraining the shape and color of the reagent bottle image where the presence of crystals results in the absence of crystals, is generated using a diffusion model and ControlNet technology to create the absence of crystals.

4. The method for generating crystal image samples based on a diffusion model according to claim 1, characterized in that, The crystal image generation, performed via a diffusion model, also includes; The LoRA model is used to constrain the shape and color conditions for generating crystal images.

5. The method for generating crystal image samples based on a diffusion model according to claim 1, characterized in that, The crystal image generation, performed via a diffusion model, also includes: The crystal image is then rendered in HDR.