Wind power cabin fire identification and positioning method and system based on diffusion model

By generating high-quality composite images of fires using diffusion models and data generation techniques, the accuracy and precision issues of wind turbine nacelle fire monitoring models have been resolved, enabling efficient and automatic identification and location of wind turbine nacelle fires.

CN121884342APending Publication Date: 2026-04-17LONGYUAN BEIJING WIND POWER ENG TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
LONGYUAN BEIJING WIND POWER ENG TECH
Filing Date
2025-11-28
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing deep learning-based fire monitoring models suffer from low accuracy and poor positioning precision in wind turbine nacelle scenarios, and lack high-quality real-world fire training data.

Method used

By using a diffusion model-based approach, the model is fine-tuned using an image-text dataset of wind turbine nacelle scene images and fire scene text descriptions. High-quality synthetic fire images are generated by combining noise inversion technology and a spatial structure control network. Quality scores are then used for screening and automatic annotation to form an annotated dataset, which is then used to train fire identification, classification, and localization models.

Benefits of technology

It improves the accuracy and positioning precision of fire identification in wind turbine nacelles, enhances the reliability of fire monitoring, and provides an efficient and low-cost safety monitoring solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121884342A_ABST
    Figure CN121884342A_ABST
Patent Text Reader

Abstract

The invention provides a wind power cabin fire identification and positioning method and system based on a diffusion model, and the method comprises the steps: obtaining a diffusion model matched with a fire scene based on a pre-trained text-to-image diffusion model through employing an image-text pair data set containing a wind power cabin scene image and a fire scene text description, in combination with a noise inversion technology and / or a spatial structure control network, generating a fire synthetic image with real wind power cabin scene features; performing quality score screening on the fire synthetic image, retaining images with scores higher than a set threshold value, and performing automatic fire area labeling on the screened images by using a vision-language model to form a labeling data set; and respectively training a fire identification and classification model and a fire area positioning model by using the labeled data set. The problem that high-quality fire training data of a real wind power cabin scene is difficult to obtain is effectively solved, and the fire recognition and positioning accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of fire location and identification technology, and in particular to a method and system for fire identification and location in wind turbine nacelles based on a diffusion model. Background Technology

[0002] The internal structure of wind turbine nacelles is complex and the environment is unique, often operating unattended for extended periods. Therefore, accurate early fire identification and location are crucial for preventing major accidents. Currently, deep learning-based fire monitoring models heavily rely on large amounts of high-quality labeled data for training. However, obtaining realistic images of wind turbine nacelle fires is extremely difficult in practice. Existing publicly available fire datasets primarily depict outdoor scenes, which differ significantly from the indoor environment of wind turbine nacelles. This results in models trained directly using these datasets exhibiting low accuracy in wind turbine nacelle scenarios, poor localization precision, and severely insufficient generalization ability.

[0003] Therefore, how to effectively obtain high-quality fire training data for wind turbine nacelle scenarios without relying on real fire data has become a technical bottleneck that urgently needs to be solved in this field. Summary of the Invention

[0004] This invention provides a method and system for identifying and locating fires in wind turbine nacelles based on a diffusion model, which addresses the shortcomings of existing technologies that struggle to obtain high-quality fire training data for real wind turbine nacelle scenarios, resulting in poor accuracy of deep learning-based fire monitoring models.

[0005] In a first aspect, the present invention provides a method for identifying and locating fires in wind turbine nacelles based on a diffusion model, comprising:

[0006] Based on a pre-trained text-to-image diffusion model, using an image-text pair dataset containing images of wind turbine nacelle scenes and text descriptions of fire scenes, a parameter-efficient fine-tuning method is employed to fine-tune the model and obtain a diffusion model adapted to fire scenes.

[0007] Using the aforementioned diffusion model, combined with noise inversion technology and / or a spatial structure control network, a composite fire image with the characteristics of a real wind turbine nacelle scene is generated.

[0008] The composite fire images are screened by quality scoring, and images with scores higher than a set threshold are retained. The fire areas of the screened images are automatically labeled using a visual-language model to form a labeled dataset.

[0009] The labeled dataset was used to train a fire identification and classification model and a fire area location model, respectively. The trained models were then deployed in the wind turbine nacelle monitoring system to achieve automatic identification and location of wind turbine nacelle fires.

[0010] According to the present invention, a wind turbine nacelle fire identification and localization method based on a diffusion model is provided, which constructs an image-text pair dataset for fine-tuning, including:

[0011] Collect images of the interior of wind turbine nacelles and various indoor fire images;

[0012] Use a visual-language model to generate detailed description text for each image;

[0013] The detailed description text was cleaned, and keywords related to fire were removed.

[0014] Preset trigger words are added before the cleaned text description to form the final text prompts for model fine-tuning.

[0015] According to the present invention, a method for identifying and locating fires in a wind turbine nacelle based on a diffusion model is provided. The method employs a parameter-efficient fine-tuning method to fine-tune the model and obtain a diffusion model adapted to the fire scenario, comprising:

[0016] A low-rank adaptation matrix is ​​injected into the attention module of the text-to-image diffusion model, the original model parameters are frozen, and only the low-rank adaptation matrix is ​​trained.

[0017] By fine-tuning the trained text-to-image diffusion model using an image-text pair dataset containing images of wind turbine nacelle scenes and text descriptions of fire scenes, a diffusion model adapted to fire scenes can be obtained.

[0018] According to the present invention, a method for identifying and locating wind turbine nacelle fires based on a diffusion model is provided, wherein generating a composite fire image with real wind turbine nacelle scene characteristics by combining a spatial structure control network includes:

[0019] Select the input image of the wind turbine nacelle scene;

[0020] Semantic segmentation or depth estimation is performed on the wind turbine nacelle scene image to extract scene structure information;

[0021] The scene structure information is input into the spatial structure control network as a control condition, and a fire scene-adapted diffusion model is used to generate a wind turbine nacelle fire image with consistent structure.

[0022] According to the present invention, a method for identifying and locating wind turbine nacelle fires based on a diffusion model, wherein the method combines noise inversion technology to generate a composite fire image with characteristics of a real wind turbine nacelle scene, comprising:

[0023] Noise is added to the input wind turbine nacelle scene image through a forward diffusion process;

[0024] During the reverse spread process, text prompts describing the fire are used as a condition;

[0025] By controlling the noise level of the back diffusion process, a synthetic image is generated that is consistent with the content of the input image and contains fire features.

[0026] According to the present invention, a method for identifying and locating wind turbine nacelle fires based on a diffusion model is provided, wherein the quality scoring and screening of the synthesized fire image includes:

[0027] Use a visual-language model to evaluate the semantic consistency between the generated image and the text prompt;

[0028] An image quality assessment model is used to score the sharpness and naturalness of the generated images;

[0029] The credibility of fire features in images is determined using a pre-trained fire recognition model.

[0030] The overall score is determined by combining the semantic consistency, clarity, and naturalness scores with the credibility score.

[0031] According to the present invention, a method for identifying and locating fires in wind turbine nacelles based on a diffusion model is provided, wherein the method automatically labels the fire areas in the filtered images using a visual-language model to form a labeled dataset, including:

[0032] Use a visual-language model to identify fire areas in filtered images and generate text descriptions;

[0033] Construct segmentation suggestions based on the text description;

[0034] The original image and the segmentation prompts are input into the prompt-based segmentation model to obtain the pixel-level fire area mask output by the segmentation model, thus completing the automatic annotation.

[0035] According to the present invention, a method for identifying and locating fires in a wind turbine nacelle based on a diffusion model is provided, wherein training a fire identification classification model and a fire area location model using the labeled dataset includes:

[0036] A fire identification and classification model based on a convolutional neural network was trained using a labeled dataset.

[0037] A fire area localization model based on an object detection network was trained using a dataset with fire area annotations.

[0038] The trained model is tested on a validation set to select the model weights with the best performance. The best model is then converted into an inference format and deployed in the wind turbine nacelle monitoring system.

[0039] The wind turbine nacelle fire identification and location method based on a diffusion model provided by the present invention further includes:

[0040] The trained fire identification and classification model and fire area location model were deployed on the server in the wind farm control room.

[0041] Configure a real-time communication link between the wind turbine nacelle monitoring camera and the server;

[0042] Establish a fire early warning mechanism that automatically triggers an alarm and locates the fire source when a fire is detected;

[0043] Design a visual interface to display monitoring footage and fire identification results in real time.

[0044] Secondly, the present invention also provides a wind turbine nacelle fire identification and location system based on a diffusion model, comprising:

[0045] The adaptation module is used to fine-tune the model based on the pre-trained text-to-image diffusion model. It utilizes an image-text pair dataset containing images of wind turbine nacelle scenes and text descriptions of fire scenes, and employs an efficient parameter fine-tuning method to obtain a diffusion model adapted to the fire scene.

[0046] A synthesis module is used to generate a composite fire image with real wind turbine nacelle scene characteristics by using the diffusion model, combined with noise inversion technology and / or a spatial structure control network.

[0047] The annotation module is used to perform quality scoring and screening on the composite fire images, retaining images with scores higher than a set threshold, and automatically annotating the fire areas of the screened images using a visual-language model to form an annotation dataset.

[0048] The training module is used to train the fire identification and classification model and the fire area location model using the labeled dataset, and then deploys the trained models to the wind turbine nacelle monitoring system to achieve automatic identification and location of wind turbine nacelle fires.

[0049] Thirdly, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the wind turbine nacelle fire identification and location method based on the diffusion model as described above.

[0050] Fourthly, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the wind turbine nacelle fire identification and location method based on the diffusion model as described above.

[0051] Fifthly, the present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the wind turbine nacelle fire identification and location method based on the diffusion model as described above.

[0052] This invention provides a method and system for identifying and locating wind turbine nacelle fires based on a diffusion model. It adapts a pre-trained diffusion model to wind turbine nacelle fire scenarios through efficient parameter fine-tuning, and generates high-quality, reliable synthetic fire images with high scenario fit by combining noise inversion and spatial structure control techniques. Furthermore, it establishes an automated data cleaning and annotation process, effectively addressing the problem of insufficient real-world fire training data for specific wind turbine nacelle scenarios. The fire identification classification model and fire area location model trained using the annotated dataset constructed using this method significantly improve the accuracy of fire identification in wind turbine nacelle fire identification tasks and the precision of fire location tasks, significantly enhancing the reliability of wind turbine nacelle fire monitoring and providing the wind power industry with an efficient and low-cost safety monitoring solution. Attached Figure Description

[0053] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0054] Figure 1 This is a flowchart illustrating the wind turbine nacelle fire identification and location method based on a diffusion model provided in this embodiment.

[0055] Figure 2 This is a schematic diagram illustrating the principle of collecting fire scenes, scene descriptions, and adding trigger words provided in this embodiment;

[0056] Figure 3 This is a schematic diagram of a wind turbine nacelle scenario provided in this embodiment;

[0057] Figure 4 This is a schematic diagram of a wind turbine nacelle fire provided in this embodiment;

[0058] Figure 5 This is a visualization diagram of the annotation results provided in this embodiment;

[0059] Figure 6 This is a schematic diagram of the structure of the wind turbine nacelle fire identification and location system based on the diffusion model provided in this embodiment;

[0060] Figure 7 This is a schematic diagram of the structure of the electronic device provided in this embodiment. Detailed Implementation

[0061] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0062] Figure 1 This is a flowchart illustrating the wind turbine nacelle fire identification and location method based on a diffusion model provided in this embodiment.

[0063] like Figure 1 As shown in the figure, the wind turbine nacelle fire identification and location method based on a diffusion model provided by this invention aims to solve the technical bottleneck of low identification accuracy and poor location accuracy in existing wind turbine nacelle fire monitoring models due to a lack of real-world training data. The method mainly includes the following steps:

[0064] 101. Based on a pre-trained text-to-image diffusion model, using an image-text pair dataset containing images of wind turbine nacelle scenes and text descriptions of fire scenes, a parameter-efficient fine-tuning method is employed to fine-tune the model and obtain a diffusion model adapted to fire scenes.

[0065] Specifically, to adapt the pre-trained diffusion model to wind turbine nacelle fire scenarios, it is necessary to construct a dedicated image-text pair dataset and employ an efficient parameter fine-tuning method, as follows:

[0066] To ensure the fine-tuned model accurately responds to the fire generation requirements of wind turbine nacelle scenarios, a scene-fitting image-text pair dataset needs to be constructed first. Images of the interior of wind turbine nacelles are collected from publicly available internet resources, wind power company equipment archives, and relevant incident records. These images cover core equipment areas such as control cabinets, gearboxes, cable trays, and ventilation ducts, including footage from different shooting angles, with varying equipment operating states, and under different lighting conditions. Simultaneously, various indoor fire images are collected, primarily focusing on electrical fires and equipment combustion fires, covering scenarios with different types of combustibles and different stages of fire development. Figure 2 As shown on the left.

[0067] A visual-language model is used to generate structured, detailed description text for each collected image, such as... Figure 2 The content in the upper right corner. For example, for an image of the control cabinet inside a wind turbine nacelle, generate a description that reflects the characteristics of the scene; for an image of an indoor equipment fire, generate a description that includes non-core fire features such as equipment type and environmental background, ensuring that the text can fully reflect the scene details of the image.

[0068] By combining manual verification with automated rules, all keywords related to "fire" in the descriptive text, such as "fire," "flame," "burning," "scorch marks," and "smoke," are removed. This avoids the model learning fire features of non-target scenarios in advance during the fine-tuning process, ensuring that the text retains only the feature information of the scene itself.

[0069] At the beginning of each cleaned text description, a pre-defined trigger word, such as "SDFOD," is uniformly added to form the final text prompts for model fine-tuning. The purpose of these trigger words is to help the model accurately identify the scene boundaries of the fine-tuning task, improving the scene fit of subsequently generated images. For example... Figure 2 As shown in the bottom right corner.

[0070] An open-source text-to-image diffusion model (such as Stable Diffusion) was chosen as the pre-trained base model. This model possesses mature high-quality image generation capabilities, and its open-source nature facilitates subsequent parameter tuning and optimization. A low-rank adaptation (LoRA) matrix was injected into the attention module (such as the Cross Attention layer) of the base model, decomposing the original weight matrix of the attention module into two low-rank matrices. This method enables efficient parameter fine-tuning. Simultaneously, all original parameters of the base model were frozen, retaining only the injected low-rank adaptation matrix as trainable parameters. This significantly reduces the computational resource consumption of the fine-tuning process and avoids compromising the scene generation capabilities of the original model.

[0071] The model was fine-tuned using the constructed image-text pair dataset described above. The training process optimized the semantic matching degree of the image-text pairs until the consistency between the images generated by the model and the text prompts on the validation set reached a stable state, ultimately yielding a diffusion model adapted to fire scenarios. This diffusion model retains the image generation quality of the base model while accurately responding to the fire generation requirements in wind turbine nacelle scenarios.

[0072] 102. Using a diffusion model, combined with noise inversion technology and / or a spatial structure control network, generate a composite fire image with the characteristics of a real wind turbine nacelle scene.

[0073] Specifically, a fine-tuned diffusion model is used, combined with noise inversion techniques and / or a spatial structure control network, to generate synthetic images, ensuring the scene realism and structural rationality of the generated images.

[0074] The image generated using noise inversion techniques is as follows:

[0075] The wind turbine nacelle scene image corresponding to the scene to be generated is selected as the input and fed into the forward diffusion process of the diffusion model adapted to the fire scene. Noise is gradually added to the input image according to the noise scheduler preset by the model until the input image is completely transformed into a noise map, thus completing the forward diffusion process.

[0076] During the backdiffusion process, a text prompt containing fire feature descriptions (the text must include wind turbine nacelle scene information and fire feature information, and begin with the trigger word "SDFOD") is used as a condition. By controlling the noise attenuation rate at each step of the backdiffusion process, the model gradually removes noise, restores image content, and integrates fire features into the wind turbine nacelle scene, ultimately generating a synthetic image that is consistent with the input wind turbine nacelle image and contains real fire features. This method ensures that the scene background of the generated image is highly consistent with the input nacelle image, preventing the fire scene from deviating from the real nacelle environment.

[0077] The image generated by the spatial structure control network is as follows:

[0078] Select a target wind turbine nacelle scene image and perform semantic segmentation on the image using a semantic segmentation model (such as SegFormer) to identify and extract semantic masks of the core equipment areas (such as control cabinets, gearboxes, cables, etc.) within the wind turbine nacelle. Alternatively, use a depth estimation model (such as DPT-Hybrid) to estimate the depth of the image and generate a depth map reflecting the spatial relationships of the equipment within the nacelle. The semantic masks and depth maps together constitute the structural information of the wind turbine nacelle scene.

[0079] The extracted structural information is input into the spatial structure control network ControlNet, and ControlNet is associated with a diffusion model adapted to the fire scene. Text prompts containing fire features are used as generation conditions. ControlNet imposes structural constraints on the generation process of the diffusion model to ensure that the fire image generated by the model is completely consistent with the input wind turbine nacelle image in terms of spatial structure. For example, the fire only appears on the surface of the equipment and does not leave the equipment area. Finally, a composite image of a wind turbine nacelle fire with reasonable structure and realistic scene is generated.

[0080] To further improve the quality of the generated images, noise inversion technology and a spatial structure control network can be combined simultaneously. During the generation process, noise inversion technology is first used to ensure that the scene content of the generated image is consistent with the input nacelle image. Then, ControlNet is used to constrain the spatial structure of the generated image. Simultaneously, text prompts containing fire features are used as conditions, enabling the model to generate a synthetic image of a wind turbine nacelle fire that combines scene realism, structural rationality, and accuracy of fire features. For example... Figure 3 The image shown is a scene of a wind turbine nacelle. Figure 4 This is a diagram of a fire in a wind turbine nacelle.

[0081] By employing three generation methods—noise-only inversion, spatial structure-only control, and a combination of both techniques—composite fire images covering different core areas and stages of fire development within the wind turbine nacelle can be generated, meeting the diverse needs of subsequent model training.

[0082] 103. The composite fire images are screened by quality scoring, and images with scores higher than a set threshold are retained. The fire areas of the screened images are automatically labeled using a vision-language model to form a labeled dataset.

[0083] Specifically, a high-quality labeled dataset is formed through quality scoring and automatic annotation to ensure that the dataset meets the model training requirements.

[0084] A visual-language model (such as Florence-2-large) is used to evaluate the semantic matching degree between the generated synthetic fire image and the corresponding text prompt. This determines whether the generated image fully presents the wind turbine nacelle scene features and fire features in the text prompt, and a semantic consistency score is obtained.

[0085] Image quality assessment models (such as NIQE) are used to score the generated images for indicators such as sharpness, color naturalness, and fire texture realism, ensuring that the generated images are free from problems such as blurriness, color distortion, and texture abnormalities, and thus obtaining sharpness and naturalness scores.

[0086] A pre-trained fire identification model (such as the ResNet series models) is used to identify fire features in the generated images, and to determine whether the fire shape, color, and burning state in the images conform to the characteristics of a real electrical fire, thereby obtaining a fire feature credibility score.

[0087] Weights for each scoring metric are set according to actual needs. The semantic consistency score, clarity and naturalness score, and fire feature credibility score are comprehensively calculated to obtain the total score for each generated image. A scoring threshold is set; images with a total score above the threshold are retained as qualified synthesized images, while low-quality images with a total score below the threshold are deleted, ensuring high-quality training data for subsequent models. Figure 4 The image shown is a qualified composite image of a fire.

[0088] Visual-language models (such as Florence-2-large and RAM) are used to identify fire areas in the selected qualified composite fire images and generate text descriptions containing the location, extent, and morphological features of the fire areas. For example, the fire area is located on the surface of the wind turbine nacelle control cabinet, has an irregular shape, and is orange-yellow in color.

[0089] Based on the generated text description of the fire area and combined with the image pixel coordinate information, segmentation prompts for image segmentation are constructed, such as segmenting the orange-yellow fire area on the surface of the wind turbine nacelle control cabinet in the image, and marking the core pixels of the fire area as auxiliary prompts to improve segmentation accuracy.

[0090] A qualified composite fire image and the constructed segmentation cue are input into a cue-based segmentation model (SAM). SAM performs pixel-level segmentation of the fire region in the image and outputs a pixel-level mask of the fire region. Each pixel belonging to the fire region is marked in the mask, thus completing the automatic labeling of the fire region. Figure 5 As shown.

[0091] The qualified fire composite images that have been automatically annotated are associated with the corresponding pixel-level masks and image description text. The dataset is then divided into training, validation and test sets according to the model training requirements to form an annotated dataset for subsequent model training.

[0092] 104. Use labeled datasets to train fire identification and classification models and fire area location models respectively, and deploy the trained models on the wind turbine nacelle monitoring system to achieve automatic identification and location of wind turbine nacelle fires.

[0093] Specifically, a model is trained and deployed using labeled datasets to achieve automatic identification and location of fires in wind turbine nacelles.

[0094] Lightweight convolutional neural networks (such as ResNet and MobileNet) were selected as the basic architecture for the fire identification and classification model. The training set of the labeled dataset was used as the training data. Synthetic fire images and original wind turbine nacelle scene images (without fire) were used as input to the model. The model output target was a binary classification result of "fire present / no fire." During training, the model performance was monitored in real time using a validation set. Training was stopped when the model's recognition accuracy on the validation set reached a stable state. The model weights with the best performance were selected, completing the training of the fire identification and classification model. This model can quickly determine whether a fire exists in a wind turbine nacelle scene, providing a prerequisite for subsequent localization.

[0095] An object detection network (Fast R-CNN, YOLO) was selected as the basic architecture of the fire area localization model. The training set (containing pixel-level fire area annotations) of the labeled dataset was used as the training data, and synthetic fire images were used as the model input. The bounding box coordinates and confidence scores of the fire area were used as the model output targets for training. The model's localization accuracy (e.g., IoU metric) was monitored using a validation set. Training was stopped when the localization accuracy reached a stable state, and the model weights with the best performance were selected to complete the training of the fire area localization model. This model can accurately define the location and extent of fire areas, providing precise data for firefighting and maintenance.

[0096] The trained fire identification and classification model and fire area location model are converted into a format suitable for real-time inference, such as ONNX format, and deployed on the server in the wind farm control room to ensure that the server has the computing power to meet the real-time inference requirements of the model.

[0097] High-temperature resistant and vibration-resistant monitoring cameras are installed inside the wind turbine nacelle. A real-time communication link is established between the monitoring cameras and the control room server via industrial Ethernet to ensure that the real-time video streams of the wind turbine nacelle captured by the cameras can be stably transmitted to the server.

[0098] The server is configured with early warning trigger conditions. When the fire identification and classification model determines that there is a fire in the video stream frame and the confidence level output by the fire area positioning model reaches the set threshold, the server automatically triggers an audible and visual alarm and pushes the fire alarm information (including the cabin number and the coordinates of the fire source) to the mobile devices of the maintenance personnel to achieve rapid fire early warning.

[0099] The algorithm logic is implemented in Python, and the real-time inference module is developed in C++, integrating them to form wind turbine nacelle disaster early warning and disaster location software. A visual interface is designed within the software to display real-time wind turbine nacelle monitoring footage, fire identification results, fire location bounding boxes, and historical alarm records, facilitating real-time monitoring and retrospective analysis by maintenance personnel. Ultimately, the monitoring cameras, control room server, and early warning software are integrated into a complete electronic device for wind turbine nacelle fire monitoring, achieving full automation from data acquisition and model inference to early warning display.

[0100] Based on the same general inventive concept, this invention also protects a wind turbine nacelle fire identification and location system based on a diffusion model. The wind turbine nacelle fire identification and location system based on a diffusion model described below and the wind turbine nacelle fire identification and location method based on a diffusion model described above can be referred to in correspondence.

[0101] Figure 6 This is a schematic diagram of the structure of the wind turbine nacelle fire identification and location system based on the diffusion model provided in this embodiment.

[0102] like Figure 6 As shown in the figure, this embodiment provides a wind turbine nacelle fire identification and location system based on a diffusion model, comprising:

[0103] The adaptation module 601 is used to fine-tune the model based on the pre-trained text-to-image diffusion model. It uses an image-text pair dataset containing images of wind turbine nacelle scenes and text descriptions of fire scenes, and employs an efficient parameter fine-tuning method to obtain a diffusion model adapted to the fire scene.

[0104] Synthesis module 602 is used to generate a composite fire image with real wind turbine nacelle scene characteristics by using a diffusion model, combined with noise inversion technology and / or a spatial structure control network.

[0105] The annotation module 603 is used to perform quality scoring and screening on the composite fire images, retaining images with scores higher than a set threshold, and automatically annotating the fire areas of the filtered images using a vision-language model to form an annotation dataset.

[0106] Training module 604 is used to train a fire identification and classification model and a fire area location model using labeled datasets, and then deploys the trained models to the wind turbine nacelle monitoring system to achieve automatic identification and location of wind turbine nacelle fires.

[0107] Figure 7 This is a schematic diagram of the structure of the electronic device provided in this embodiment.

[0108] like Figure 7 As shown, the electronic device may include a processor 701, a communications interface 702, a memory 703, and a communication bus 704. The processor 701, communications interface 702, and memory 703 communicate with each other via the communication bus 704. The processor 701 can call logical instructions from the memory 703 to execute a diffusion-based wind turbine nacelle fire identification and location method.

[0109] Furthermore, the logical instructions in the aforementioned memory 703 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0110] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the wind turbine nacelle fire identification and location method based on the diffusion model provided by the above methods.

[0111] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the diffusion model-based wind turbine nacelle fire identification and location method provided by the above methods.

[0112] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0113] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0114] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for identifying and locating fires in wind turbine nacelles based on a diffusion model, characterized in that, include: Based on a pre-trained text-to-image diffusion model, using an image-text pair dataset containing images of wind turbine nacelle scenes and text descriptions of fire scenes, a parameter-efficient fine-tuning method is employed to fine-tune the model and obtain a diffusion model adapted to fire scenes. Using the aforementioned diffusion model, combined with noise inversion technology and / or a spatial structure control network, a composite fire image with the characteristics of a real wind turbine nacelle scene is generated. The composite fire images are screened by quality scoring, and images with scores higher than a set threshold are retained. The fire areas of the screened images are automatically labeled using a visual-language model to form a labeled dataset. The labeled dataset was used to train a fire identification and classification model and a fire area location model, respectively. The trained models were then deployed in the wind turbine nacelle monitoring system to achieve automatic identification and location of wind turbine nacelle fires.

2. The wind turbine nacelle fire identification and location method based on a diffusion model according to claim 1, characterized in that, Construct an image-text pair dataset for fine-tuning, including: Collect images of the interior of wind turbine nacelles and various indoor fire images; Use a visual-language model to generate detailed description text for each image; The detailed description text was cleaned, and keywords related to fire were removed. Preset trigger words are added before the cleaned text description to form the final text prompts for model fine-tuning.

3. The wind turbine nacelle fire identification and location method based on a diffusion model according to claim 1, characterized in that, The method of using efficient parameter fine-tuning to fine-tune the model and obtain a diffusion model adapted to the fire scenario includes: A low-rank adaptation matrix is ​​injected into the attention module of the text-to-image diffusion model, the original model parameters are frozen, and only the low-rank adaptation matrix is ​​trained. By fine-tuning the trained text-to-image diffusion model using an image-text pair dataset containing images of wind turbine nacelle scenes and text descriptions of fire scenes, a diffusion model adapted to fire scenes can be obtained.

4. The wind turbine nacelle fire identification and location method based on a diffusion model according to claim 1, characterized in that, The method of generating a composite fire image with realistic wind turbine nacelle scene characteristics by combining a spatial structure control network includes: Select the input image of the wind turbine nacelle scene; Semantic segmentation or depth estimation is performed on the wind turbine nacelle scene image to extract scene structure information; The scene structure information is input into the spatial structure control network as a control condition, and a fire scene-adapted diffusion model is used to generate a wind turbine nacelle fire image with consistent structure.

5. The wind turbine nacelle fire identification and location method based on a diffusion model according to claim 1, characterized in that, The method of combining noise inversion technology to generate a composite fire image with characteristics of a real wind turbine nacelle scene includes: Noise is added to the input wind turbine nacelle scene image through a forward diffusion process; During the reverse spread process, text prompts describing the fire are used as a condition; By controlling the noise level of the back diffusion process, a synthetic image is generated that is consistent with the content of the input image and contains fire features.

6. The wind turbine nacelle fire identification and location method based on a diffusion model according to claim 1, characterized in that, The quality scoring and screening of the composite fire image includes: Use a visual-language model to evaluate the semantic consistency between the generated image and the text prompt; An image quality assessment model is used to score the sharpness and naturalness of the generated images; The credibility of fire features in images is determined using a pre-trained fire recognition model. The overall score is determined by combining the semantic consistency, clarity, and naturalness scores with the credibility score.

7. The wind turbine nacelle fire identification and location method based on a diffusion model according to claim 1, characterized in that, The process of automatically labeling fire areas in the selected images using a vision-language model to form a labeled dataset includes: Use a visual-language model to identify fire areas in filtered images and generate text descriptions; Construct segmentation suggestions based on the text description; The original image and the segmentation prompts are input into the prompt-based segmentation model to obtain the pixel-level fire area mask output by the segmentation model, thus completing the automatic annotation.

8. The wind turbine nacelle fire identification and location method based on a diffusion model according to claim 1, characterized in that, The step of training the fire identification and classification model and the fire area localization model using the labeled dataset includes: A fire identification and classification model based on a convolutional neural network was trained using a labeled dataset. A fire area localization model based on an object detection network was trained using a dataset with fire area annotations. The trained model is tested on a validation set to select the model weights with the best performance. The best model is then converted into an inference format and deployed in the wind turbine nacelle monitoring system.

9. The wind turbine nacelle fire identification and location method based on a diffusion model according to claim 8, characterized in that, Also includes: The trained fire identification and classification model and fire area location model were deployed on the server in the wind farm control room. Configure a real-time communication link between the wind turbine nacelle monitoring camera and the server; Establish a fire early warning mechanism that automatically triggers an alarm and locates the fire source when a fire is detected; Design a visual interface to display monitoring footage and fire identification results in real time.

10. A wind turbine nacelle fire identification and location system based on a diffusion model, characterized in that, include: The adaptation module is used to fine-tune the model based on the pre-trained text-to-image diffusion model. It utilizes an image-text pair dataset containing images of wind turbine nacelle scenes and text descriptions of fire scenes, and employs an efficient parameter fine-tuning method to obtain a diffusion model adapted to the fire scene. A synthesis module is used to generate a composite fire image with real wind turbine nacelle scene characteristics by using the diffusion model, combined with noise inversion technology and / or a spatial structure control network. The annotation module is used to perform quality scoring and screening on the composite fire images, retaining images with scores higher than a set threshold, and automatically annotating the fire areas of the screened images using a visual-language model to form an annotation dataset. The training module is used to train the fire identification and classification model and the fire area location model using the labeled dataset, and then deploys the trained models to the wind turbine nacelle monitoring system to achieve automatic identification and location of wind turbine nacelle fires.