Safety education image generation method based on stable diffusion and LiblibAI

By combining a stable diffusion model with the LiblibAI platform, this study solves the technical problems that existing technologies have failed to address in generating safety education images. It achieves efficient, realistic, and controllable generation of these images.

CN121120852APending Publication Date: 2025-12-12SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511021589.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-24
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

Existing technologies for generating safety education images suffer from problems such as distortion of accident elements, misalignment of equipment and scenes, and difficulty in accurately controlling the morphological details of traditional prompts. As a result, the generated images cannot meet the professional requirements of the engineering field for high fidelity and strong scene-based representation.

Method used

By combining a stable diffusion model with the LiblibAI platform, and through structured prompt word design, graph-to-graph reconstruction, and ControlNet edge constraint technology, a complete workflow from data preprocessing to image output is constructed to achieve realistic reproduction of dangerous scenarios such as fires and explosions.

Benefits of technology

It achieves efficient and controllable generation of safety education images, improves image generation efficiency, enhances realism and visual impact, supports semantic control and contextual consistency of image content, and is suitable for vocational education and safety drills.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121120852A_ABST
    Figure CN121120852A_ABST
Patent Text Reader

Abstract

The invention discloses a safety education image generation method based on stable diffusion and LiblibAI. The method comprises the following steps: collecting and preprocessing image data related to a safety accident; constructing a structured cue word library through stable diffusion reverse cue extraction and hierarchical optimization; a basic model and a LoRA model (a sea corner FLUX flame element V1 and a sea corner Flux shadow art V1.0) are integrated on a LiblibAI platform, and physical simulation of flames and shadows is achieved through parameter collaborative tuning; based on a real equipment photo, an accident scene is generated by combining an image-to-image generation mode with a ControlNet edge guidance technology; and outputting a safety education image resource library meeting teaching requirements. The problems that a traditional safety teaching material image is insufficient in authenticity and high in updating cost are solved, and high-fidelity and strong-scene safety education visual resource production is achieved through the controllable AI generation technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence education application technology, specifically involving a method for generating safety education images based on stable diffusion and the LiblibAI platform, which is suitable for the development of visualization resources for professional education scenarios such as laboratory safety and industrial safety. Background Technology

[0002] Safety education plays a crucial role in the engineering field, but traditional safety teaching images are mainly created through photography, hand-drawing, or computer graphics. These methods have significant limitations: high production costs, with single-scene construction costs reaching thousands of yuan or even higher; long production cycles, with complex accident scenarios requiring weeks of manual drawing; and difficulties in updating and maintaining, necessitating the creation of entirely new materials for new equipment or specific accident types. These bottlenecks result in a severe shortage of teaching resources, making it difficult to meet the urgent need for high-quality, diverse accident simulation images in safety education.

[0003] In recent years, generative artificial intelligence technology has developed rapidly, providing a new technological path for educational image creation. International research shows that generative methods based on generative adversarial networks and diffusion models can learn specific styles through data training and output high-quality scene images. In domestic practice, open-source tools, represented by stable diffusion, have validated their generative capabilities in fields such as artistic creation and advertising design, and some research has begun to explore their application potential in professional fields such as medical image compositing and architectural visualization.

[0004] However, the application of existing technologies in safety education still faces significant challenges. Accident elements generated by general models often exhibit physical inconsistencies; for example, flames display unnatural spherical aggregate shapes, and smoke diffusion lacks fluid dynamics characteristics. Furthermore, existing methods lack targeted optimization for specialized scenarios such as equipment explosions, electrostatic discharge, and high-temperature burns, leading to spatial misalignment between equipment structures and accident elements. More critically, traditional warning image engineering struggles to precisely control morphological details and accurately simulate microscopic physical processes such as metal melting phase transitions and electric arc trajectories, severely limiting the pedagogical applicability of generated images. These technological shortcomings make current safety education image generation insufficient to meet the engineering field's professional demands for high fidelity and highly contextualized scenarios. Summary of the Invention

[0005] The purpose of this invention is to provide a high-fidelity and highly controllable method for generating safety education images. By using structured prompt word engineering, physically guided diffusion models, and edge constraint technology, it solves the core problems of accident element distortion and equipment scene misalignment in existing AI generation technologies, and achieves controllable synthesis of teaching-grade safety scenes.

[0006] The technical solution of this invention is as follows:

[0007] This invention proposes an automatic generation method for safety education images based on a stable diffusion model and the LiblibAI platform. It aims to address the problems of low efficiency, slow updates, poor realism, and lack of contextual engagement in traditional safety teaching image production. By combining an AI-generated model with advanced technologies such as structured prompt word design, graph-to-graph reconstruction, and structural edge control, this invention can realistically reproduce dangerous scenarios such as fires and explosions, and endows the images with a high degree of customization, controllability, and pedagogical applicability, providing a novel solution for vocational education, safety drills, and laboratory training.

[0008] The core of this invention lies in constructing a complete workflow from data preprocessing, semantic design, model invocation to structural control and image output. It includes the following steps:

[0009] S1: Collect real equipment images and safety accident scene data, and establish labeled datasets according to four types of accidents: fire, explosion, electric shock, and mechanical injury;

[0010] S2: Extract the original image description through the stable diffusion inverse inference module and construct a hierarchical prompt word structure;

[0011] S3: The BasicAlgorithm_F.1 model is selected on the LiblibAI platform, and the FLUX Fire ElementV1 and Flux Shadow Art V1.0 dual LoRA models are integrated;

[0012] S4: Configure dynamic parameter groups: including prompt guidance strength, noise reduction strength, and sampling steps;

[0013] S5: Upload the original image of the device and draw the outline of the accident elements in the target area using the local redraw mode;

[0014] S6: Generates physical simulation accident images using ControlNet edge-guided technology, including:

[0015] S6-1: Canny edge detection;

[0016] S6-2: Sobel operator gradient calculation;

[0017] S6-3: ControlNet parameter control;

[0018] S7. Output a series of teaching resource libraries, including static images and videos of the accident process.

[0019] At the data level, this invention first collects a large number of image samples from textbooks, safety accident case databases, and actual photos of experimental equipment to construct a preliminary labeling system and classify the image content to support the construction of prompt words and the expansion of the training set. The construction of prompt words follows a structured principle. By analyzing the response characteristics of a stable diffusion model under different input semantics, a set of contextual prompt templates with a hierarchical structure is designed. This template includes image quality prompt words (such as "high definition, UHD, 8K, ultra-fine, stereo lighting, etc.") to ensure output accuracy; style and artistic form (such as "realistic, cinematic lighting, photorealistic feel, etc.") to enhance visual realism; subject and core elements (such as "fire, flame, burning, intense heat, hell, embers, ignition, etc.") to strengthen key features through weight marking (fire: 1.5); and physical detail supplements (such as "sharp flame edges, layers of smoke, etc.") to describe microscopic physical characteristics. Among them, key elements are strengthened through weight adjustment formulas: (element name: 1.2-1.8). For example, using prompts such as "(fire:1.5), burning metal, thick smoke, realistic lighting" can effectively guide the model to generate fire scene images that meet the teaching objectives.

[0020] Regarding model selection, this invention utilizes a stable diffusion model with high robustness and good compatibility on the LiblibAI platform: BasicAlgorithm_F.1.safetensors (BasicAlgorithm_F.1 basic model). It combines the LoRA modules Cape|FLUX Fire Element V1 (weight 0.8) and Cape|Flux Shadow Art V1.0 (weight 0.8) to enhance the model's performance in flame details and shadow rendering. Cape|FLUX Fire Element V1 focuses on the physical characteristics of flames, achieving control over flame saturation and shape; Cape|Flux Shadow Art V1.0 enhances the depth of the accident scene, generating dynamic lighting and smoke layers. During image generation, the text-image guidance mechanism is implemented using the CFG scale. Higher CFG values ​​result in images that closely resemble the CFG but may reduce naturalness. Therefore, this invention selects a balance point through parameter tuning experiments, and the parameter optimization process is constrained by the latent variable update mechanism of the diffusion model. The core principle of the generation function is:

[0021] z t-1 =z t +s·(∈ cond -∈ uncond )

[0022] Among them, z t Represents the current latent variable, ∈ cond ,∈ uncondThese represent conditional / unconditional noise prediction, respectively. s is the CFG scaling factor. By controlling the parameters in this formula, this invention can achieve consistency optimization of image semantics.

[0023] To enhance the contextual consistency and device recognition of images, this invention further introduces an image-to-image method. After importing the original scene image, the method combines user-described mask regions with target prompts to regenerate and replace specific areas of the image. This approach effectively overlays dynamic elements such as flames and explosions while maintaining the background device structure, enhancing the image's credibility and readability for instruction. Specifically, mask parameters are dynamically configured: mask blur is set according to device material differences: 10-15 for metal surfaces, 20-30 for plastic surfaces; the mask fill range is fixed at 10-15 pixels. Sampling engine configuration: The Euler sampler is used, with 30-50 steps; denoising strength of 0.95 ensures detail reconstruction. Multimodal output: Static accident images and continuous video streams are generated simultaneously, with a video frame rate of 25fps supporting slow-motion playback analysis. In this stage, parameters such as denoising strength, mask blur, and sampling methods (e.g., DPM++2M) can be flexibly set to ensure a good balance between clarity and realism in the generated image.

[0024] This invention particularly emphasizes the importance of structural control. To address the insufficient accuracy of traditional texturing image methods in terms of image structure, this invention employs ControlNet as a skeleton structure guidance mechanism, introducing edge images as structural input to improve the controllability of flame and explosion patterns. Edge image generation uses the Canny algorithm (resolution 512, low threshold 100 / high threshold 200) to extract flame / explosion contours, followed by Gaussian filtering and double threshold segmentation to eliminate noise interference; the horizontal / vertical gradient G is calculated using the Sobel operator. x G y The gradient magnitude field is generated, and the core gradient calculation formula is as follows:

[0025]

[0026] Among them, G x With G yG represents the gradient changes of the image in the x and y directions, respectively, reflecting the intensity and direction information of the image edges. By performing Canny edge detection on real fire images, the extracted structure map is used as the guiding input for ControlNet. The LiblibAI platform allows setting the starting and ending control steps to control the degree of structure guidance, ensuring that the image follows structural constraints while retaining the stylistic freedom of the generative model, thus achieving a dual guarantee of realism and structural accuracy. Specifically, a parameter set (control weights 0.6-0.7, starting control step 0.1, ending control step 0.8) constrains the diffusion process, ensuring that the generated flame strictly follows the geometric topology of the input edges.

[0027] Ultimately, the teaching images generated by this invention can be organized into image sequences and keyframe sets as needed, and embedded into courseware, videos, or multimedia teaching materials to simulate fire evolution processes, explosion reaction mechanisms, etc., possessing extremely high teaching value and intuitiveness. Because the generation system supports customized prompts for any new device and scenario, it has good versatility and scalability, and can be widely applied in multiple fields such as education, industrial training, and safety drills.

[0028] The overall technical solution of this invention has significant advantages such as high degree of automation, excellent image generation quality, strong semantic control, and controllable image structure precision. It breaks through the efficiency bottleneck and content limitation of traditional teaching image generation, marking a major leap from static shooting to intelligent synthesis in the method of safety education image construction.

[0029] The beneficial effects of this invention are as follows:

[0030] 1. Significantly improves the efficiency of generating safe teaching images. This invention, by integrating a stable diffusion model, a prompt-guided mechanism, and image reconstruction technology, achieves an automated process from semantic input to teaching image output. Compared with traditional methods that rely on photography, post-editing, and hand-drawn illustrations, it greatly improves image production efficiency, enabling the generation of a teaching scene image within seconds to minutes. This significantly reduces labor costs and time consumption, making it suitable for educational institutions and enterprises that produce large quantities of image resources.

[0031] 2. Enhance the realism and visual impact of teaching images. Utilizing diffusion-based generation capabilities and LoRA effects modules such as flames and smoke, this invention can generate highly realistic and detailed fire and explosion scene images with strong visual tension and immersion, effectively stimulating learners' attention and enhancing their immersive experience and contextual understanding in teaching scenarios, significantly outperforming traditional flat illustrations and simplified diagrams.

[0032] 3. Achieve semantically controllable and targeted generation of image content. The structured prompt word design and multi-level semantic control strategy proposed in this invention can accurately specify key elements in the image, such as the location of the fire source, the intensity of the explosion, the smoke concentration, and the environmental background. At the same time, the salience of each element in the image is adjusted through weight allocation to ensure that the generated result meets the teaching intention and course requirements, realizing a teaching image construction mode of "on-demand generation" and "semantic drive".

[0033] 4. Supports image-to-image transformation, enhancing image context consistency. Through the image-to-image function in the LiblibAI platform, this invention can intelligently synthesize and redraw selected areas under fire or hazardous conditions while maintaining the original equipment or environmental image structure. The generated results retain the recognizability of the real equipment and incorporate visual changes in the disaster context, making it particularly suitable for customized teaching images based on specific spaces such as laboratories, factories, and warehouses. Attached Figure Description

[0034] Figure 1 The results of image generation based on structure control guidance are shown. In the figure, (a) is the edge map extracted from the original input image using the Canny operator, (b) is the fire image generated under the guidance of structured prompts and ControlNet edge control, the flame shape is controlled and highly consistent with the original edge structure, and (c) is the image generated without ControlNet, the flame shape obviously shows shape drift and contour instability, which shows the significant role of structure control mechanism in improving the accuracy of image structure.

[0035] Figure 2 The invention demonstrates a sequence of keyframes from a fire video, simulating the entire combustion process of a ceramic 3D printing device. Figures (a)-(d) are snapshots showing the flames gradually expanding from the initial ignition stage to the full combustion stage, respectively.

[0036] Figure 3 The invention also demonstrates a sequence of keyframes from a fire video, simulating the entire process of combustion in a metal 3D printing device. Figures (a)-(d) are snapshots of the flames gradually expanding from the initial ignition stage to the full combustion stage, respectively. Detailed Implementation

[0037] The present invention will be further illustrated below with reference to the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are for illustrative purposes only and are not intended to limit the scope of the invention.

[0038] This invention relies on the collaborative configuration of a stable diffusion model (Stable Diffusion) and the open-source image generation platform LiblibAI to achieve intelligent generation of image content for safety education scenarios. In practical applications, users first select a suitable base model as the image generation engine through the LiblibAI platform, such as the base algorithm _F.1.safetensors model. This model has strong image detail restoration and light and shadow expression capabilities, which can meet the generation needs of complex visual elements such as fire and smoke in safety education. To further enhance the special effects of the generated images, users can load LoRA models specifically for flame simulation, such as Cape|FLUX Flame Element V_1 and Cape|Flux Shadow Art V_1.0 for enhancing shadow texture. These models can work together with the main model to refine the flame dynamics, light and shadow changes, and the texture of the burning area, thereby enhancing the visual tension and realism of the generated images.

[0039] In the image generation process, the design of prompts plays a crucial guiding role. Users need to construct structured prompt phrases that include subject descriptions, fire characteristics, environmental elements, and stylistic modifications based on the required context, such as "realistic image of an art 3D printer on fire, thick smoke, flames emerging from the upper left corner, movie lights, (flames: 1.5), (smoke stream: 1.3)". The weight settings in parentheses can guide the model to enhance the expression of specific semantic components. For example, by setting the weight of the phrase "flames" to 1.5, the visibility and morphological expressiveness of the flames in the image can be significantly enhanced. The organization order of the prompts, the number of keywords, and semantic relevance all have a direct impact on the final image, and therefore require fine-tuning according to the teaching objectives.

[0040] Regarding generation parameters, the system recommends users set the sampling algorithm to DPM++2M Karras for smoother image details and structural transitions. The number of sampling steps is generally set between 25 and 35 steps to balance generation speed and image accuracy. The CFG Scale parameter controls the guiding strength of the prompts, typically set between 8 and 12, to ensure that the generated image closely follows the input semantics while maintaining creative freedom. The image resolution is recommended to be set to 768×768 or 960×640 pixels to suit the typical needs of textbook layout or courseware playback.

[0041] If users wish to directly construct a fire scenario on real equipment images, they can enable the image-to-image (img2img) mode. In this mode, users upload photos of the equipment at the scene and can manually outline the flame overlay area using a masking tool. Combined with prompts, the model is guided to naturally blend flames and smoke while preserving the original image structure and background. It is recommended to set the noise reduction intensity between 0.45 and 0.65 during this process to balance the preservation of the original image structure with the generation of new content, ensuring a seamless integration of the fire effect with the actual equipment.

[0042] Furthermore, to further improve the controllability of image structure, this invention introduces the ControlNet module as an image morphology control mechanism. Users can extract edge maps from the original image using the Canny algorithm and upload them to the platform as structural input. The ControlNet control model is set to control_sd15_canny_fp16, and with the adjustment of the starting and ending control step sizes (e.g., 0.2 and 0.9), it can ensure that key structures such as flame outlines and propagation paths remain consistent with the edge template set by the user. The control weights are generally set between 0.8 and 1.0 to ensure the effectiveness of structural guidance, allowing the image generation process to maintain a certain degree of creative freedom while following the structural input.

[0043] Once multiple images are generated in batches according to the needs of the combustion development process, users can organize the keyframe images into a coherent video sequence for classroom demonstrations of fire evolution, evacuation response, or identification of flammable parts of equipment. Using the method described in this invention, even without any image design experience, teachers or textbook writers can quickly generate high-quality, highly adaptable, and context-specific teaching image resources. The entire process requires no programming knowledge, is entirely based on a graphical user interface, and possesses excellent versatility and promotional value.

[0044] The specific implementation of this invention fully demonstrates the operability and high scalability of the AI ​​generation model in educational image generation. It has technical advantages such as mature process, controllable parameters, and strong content customization capabilities, providing a brand-new intelligent alternative to the traditional teaching image production mode.

Claims

1. A method for generating safety education images based on stable diffusion and LiblibAI, characterized in that, Includes the following steps: S1: Collect real equipment images and safety accident scene data, and establish labeled datasets according to four types of accidents: fire, explosion, electric shock, and mechanical injury; S2: Extract the original image description through the stable diffusion inverse inference module and construct a hierarchical prompt word structure; S3: On the LiblibAI platform, the basic algorithm _F.1.safetensors model is selected, and the Cape|FLUX Flame Element V_1 (weight 0.8) and Cape|Flux Shadow Art V_1.0 (weight 0.8) dual LoRA models are integrated; S4: Configure dynamic parameter groups: including prompt guidance strength, noise reduction strength, and sampling steps; S5: Upload the original image of the device and draw the outline of the accident elements in the target area using the local redraw mode; S6: Generates physical simulation accident images using ControlNet edge-guided technology, including: S6-1: Canny edge detection; S6-2: Sobel operator gradient calculation; S6-3: ControlNet parameter control; S7. Output a series of teaching resource libraries, including static images and videos of the accident process.

2. The method according to claim 1, characterized in that, The hierarchical prompt word structure in step S2 includes: S2-1: Picture quality keywords: High Definition, UHD, 8K, Ultra-fine, Stereo lighting; S2-2: Style and Art Form: Realism, Cinematic Lighting, Photorealistic Sense; S2-3: Main and core elements: fire, flame, burning, heat, hell, embers, ignition; S2-4: Physical details supplement: sharp flame edges, layers of smoke; The key elements are reinforced through a weighted adjustment formula: (Element names: 1.2-1.8).

3. The method according to claim 1, characterized in that, The model architecture in step S3 uses the basic algorithm _F.1.safetensors as the base model and integrates two LoRA modules: S3-1: Cape | FLUX Fire Element V_1 focuses on the physical properties of flames, achieving control over flame saturation and shape; S3-1: Cape | Flux Shadow Art V_1.0 enhances the depth of accident scenes and generates dynamic lighting and smoke layers; The dual models are output through a feature fusion layer, with each model having a weight of 0.8; the parameter tuning process is constrained by the latent variable update mechanism of the diffusion model. z t-1 =z t +s·(∈ cond -∈ uncond ) Where z t Let represent the current latent variable, and s be the cue guidance strength, i.e., the CFG proportionality coefficient, with a value of 3.5-7.0, ∈ cond ,∈ uncond These represent conditional and unconditional noise predictions, respectively.

4. The method according to claim 1, characterized in that, The partial redrawing operation in step S5 includes: S5-1: Dynamic configuration of mask parameters: Mask blur is set according to the differences in device material: 10-15 for metal surfaces, 20-30 for plastic surfaces; Mask fill range is fixed at 10-15 pixels. S5-2: Sampling engine configuration: uses Euler sampler, 30-50 steps; noise reduction intensity of 0.95 to ensure detail reconstruction; S5-3: Multimodal output: Simultaneously generates static accident images and continuous video streams, with a video frame rate of 25fps supporting slow-motion playback analysis.

5. The method according to claim 1, characterized in that, The ControlNet edge-guided technology in step S6 includes a three-level processing flow: S6-1: Edge feature extraction: The Canny operator is used to extract the flame / explosion contour, and Gaussian filtering and double threshold segmentation are used to eliminate noise interference. The parameters of the Canny operator are: resolution 512, low threshold 100, and high threshold 200. S6-2: Gradient Field Construction: Calculating Horizontal / Vertical Gradients G using the Sobel Operator x G y Generate gradient magnitude field: Among them, G x With G y These represent the gradient changes of the image in the x and y directions, respectively, while G reflects the intensity and direction information of the image edges. S6-3: Morphological Constraint Generation: Input the edge map into ControlNet and constrain the diffusion process through parameter sets to ensure that the generated flame strictly follows the geometric topology of the input edge. The parameter sets are control weights of 0.6-0.7, starting control step of 0.1, and ending control step of 0.8.