Substation bird nest hidden danger identification method and system based on multi-model cooperation

Through the multi-model collaboration method, combined with lightweight and strong image understanding models, quickly locate and deeply analyze the substation bird's nest, solving the problem of misidentification in complex environments and improving the accuracy and safety of substation hidden danger identification.

CN120495941APending Publication Date: 2025-08-15GUANGDONG ELECTRIC POWER SCI RES INST ENERGY TECH CO LTD

Patent Information

Application Number
CN202510680677.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The prior art is difficult to effectively distinguish non-bird's nest objects with similar colors and shapes in complex environments of substations, resulting in a high misidentification rate and affecting the accuracy of substation hidden danger identification.

Method used

A multi-model collaboration method is adopted, combining a lightweight target area recognition model and a strong image understanding image analysis model, images are collected by drones, and potential bird's nest targets are quickly positioned and tailored using small models to generate prompt word text, and deep analysis is used for large models to accurately judge the location of the bird's nest and issue alarm signals.

Benefits of technology

It improves the accuracy of the identification of the bird's nest in the substation, reduces misjudgment and misjudgment, and ensures the safe operation of substation equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495941A_ABST
    Figure CN120495941A_ABST
Patent Text Reader

Abstract

The invention discloses a transformer substation bird nest hidden danger identification method and system based on multi-model collaboration, and belongs to the technical field of computers, and the method comprises the steps: carrying out the image collection in a transformer substation scene through employing unmanned plane equipment, and obtaining a transformer substation bird nest hidden danger image set; inputting the substation bird nest hidden danger image set into a target area identification model for potential target identification and image clipping, and outputting a to-be-detected image; based on a preset recognition task, performing visual feature enhancement on the to-be-detected image, and generating a prompt word text; inputting the to-be-detected image and the cue word text into an image analysis model for bird nest recognition to obtain a recognition result; and outputting a bird nest position and an alarm signal in the to-be-detected image according to the identification result. Therefore, by implementing the method and the device, the problem that non-nest objects with similar colors and forms cannot be effectively distinguished in a complex environment in the prior art, so that the false recognition rate of hidden danger recognition of the transformer substation is high can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of computer technology, and specifically relates to a method and system for identifying bird nest hazards in substations based on multi-model collaboration. Background Art

[0002] Materials like hay and branches found in bird nests can easily cause short circuits or ground faults, leading to equipment damage, power outages, and even grid failures. This not only increases maintenance costs but also impacts power quality for users. Therefore, identifying bird nests in substations is crucial to ensuring the safe operation of power systems. However, because bird nests in substations are small, appear in many locations, and have widely varying shapes, they are difficult to accurately identify.

[0003] In substation drone inspections, bird nests are often identified using small models such as YOLO, as they typically occupy a large area and have relatively stable features. However, these small models struggle to accurately capture the core features of bird nests in complex scenarios. They are unable to effectively distinguish non-nest objects with similar colors and shapes, such as rust, and have difficulty identifying interfering elements such as grassy backgrounds within equipment cutouts. Consequently, these non-nest objects are misidentified as nests, resulting in numerous false positives and severely impacting the reliability of the recognition results. Summary of the Invention

[0004] This application proposes a method and system for identifying bird nest hazards in substations based on multi-model collaboration, which can solve the problem in the existing technology that non-bird nest objects with similar colors and shapes cannot be effectively distinguished in complex environments, resulting in a high misidentification rate in substation hazard identification.

[0005] A first aspect of the present application provides a method for identifying bird nest hazards in substations based on multi-model collaboration, the method comprising:

[0006] Use drone equipment to collect images in substation scenes and obtain a set of images of bird nest hazards in substations;

[0007] Input the substation bird nest hidden danger image set into a preset target area recognition model to perform potential target recognition and image cropping, and output an image to be detected containing potential bird nest targets;

[0008] Based on a preset recognition task, visual features of the image to be detected are enhanced to generate prompt word text;

[0009] Inputting the image to be detected and the prompt word text into a preset image analysis model to perform bird nest recognition, and obtaining a recognition result of the image to be detected;

[0010] According to the recognition result, the location of the bird's nest in the image to be detected and the corresponding alarm signal are output.

[0011] The proposed solution first uses a drone to capture substation equipment from multiple viewing angles within a substation scenario, generating a set of images identifying potential bird nest hazards. A target region recognition model is then used to locate potential bird nests within the captured images. The lightweight nature of the target region recognition model allows for rapid extraction of suspected bird nest areas. Areas irrelevant to nest detection are then removed to eliminate interference, resulting in a target image containing potential bird nests. An image analysis model capable of handling complex scenes is then used to perform a deep analysis of the target images. Visual features of non-nest objects with similar color and shape within the image are then distinguished, filtering out false positives and accurately determining the presence of bird nests on substation equipment. Each target image is then treated as a separate scene and a corresponding recognition result is output. Finally, based on the obtained recognition results, the bird nest is accurately located and an alarm signal is sent, significantly improving the accuracy of substation hazard identification and the safety of substation operations.

[0012] In a possible implementation method of the first aspect, the substation bird nest hidden danger image set is input into a preset target area recognition model to perform potential target recognition and image cropping, and an image to be detected containing potential bird nest targets is output, specifically:

[0013] Identifying potential bird nest targets from the substation bird nest hidden danger image set using the target area recognition model, and marking each potential bird nest target with a rectangular frame;

[0014] Enlarging the rectangular frame at a preset ratio, and taking the maximum value of the width and height of the enlarged rectangular frame as the side length to generate a corresponding square frame;

[0015] With the goal of retaining only the content within the square frame, the substation bird nest hidden danger image set is cropped to output the image to be detected.

[0016] This solution uses a lightweight target area recognition model to initially locate potential bird nests within a collection of substation bird nest hazard images. It identifies areas suspected of containing bird nests, then crops and magnifies these areas, providing data support for further in-depth analysis to determine if they are nests. The lightweight nature of the target area recognition model allows for rapid identification of potential bird nests, significantly improving data processing efficiency.

[0017] In a possible implementation method of the first aspect, identifying potential bird nest targets from the substation bird nest hazard image set using the target area recognition model is specifically as follows:

[0018] Extracting low-level image features from the substation bird nest hidden danger image set and generating a corresponding feature map; wherein the low-level image features include edges and textures of objects in the image;

[0019] By performing feature splicing, upsampling and feature fusion on the feature map, potential bird nest targets in the substation bird nest hidden danger image set are identified.

[0020] The above solution extracts the shape features of objects from the image set of bird nest hidden dangers in substations, and quickly identifies areas where bird nests are suspected by splicing, upsampling and fusing the features.

[0021] In a possible implementation method of the first aspect, before inputting the substation bird nest hazard image set into a preset target area recognition model, the method further includes:

[0022] Performing image enhancement, horizontal flipping, image transformation, visual adjustment, and target copying on the substation bird nest hidden danger image set in sequence;

[0023] Among them, the image enhancement is to stitch several images into one image; the image transformation is to translate and scale the image; the visual adjustment is to adjust the hue, saturation and brightness of the image; the target copying is to copy and paste the target in other images to the current image.

[0024] This solution improves image quality by preprocessing the collected images of bird nest hazards at substations. Image enhancement can simulate complex scenes and enhance multi-scale object detection. Visual adjustments improve image quality, while object replication addresses the long-tail distribution of bird nest morphology and reduces the risk of uneven data distribution.

[0025] In a possible implementation method of the first aspect, based on a preset recognition task, visual features of the image to be detected are enhanced to generate prompt word text, specifically:

[0026] Based on a preset recognition task, the dimension judgment and visual feature enhancement of the potential bird's nest target in the image to be detected are performed, and corresponding prompt text is generated for each image to be detected; wherein, the recognition task includes identifying hay and identifying rust.

[0027] This solution extracts text from the image to be inspected and generates prompt text for the image analysis model to identify features. Because the image analysis model has strong image understanding capabilities and can handle complex scenarios, it can accurately identify misidentified scenes in each image to be inspected. By analyzing visual features such as the structure and material of potential bird nests, the prompt text helps quickly determine whether a potential bird nest in the image is a nest and whether it is located on substation equipment.

[0028] In a possible implementation method of the first aspect, based on a preset recognition task, dimension judgment and visual feature enhancement are performed on the potential bird's nest target in the image to be detected, and corresponding prompt text is generated for each image to be detected, specifically:

[0029] Constructing a judgment matrix based on the spatial relationship between the identification target of the identification task and the substation equipment, and the existence of the identification target;

[0030] According to the recognition task, the color features, material features and morphological features of the potential bird's nest target in the image to be detected are enhanced;

[0031] The image to be detected after feature enhancement is combined with the judgment matrix to generate the prompt text.

[0032] This approach constructs a judgment matrix based on the presence of the target in the image and the spatial relationship between the target and substation equipment. It then enhances the visual features of potential bird nests in the image, including color, material, and morphology. This eliminates interference for subsequent nest identification and creates a clear visual contrast with misidentified nests, facilitating accurate subsequent identification.

[0033] In a possible implementation method of the first aspect, the image to be detected and the prompt word text are input into a preset image analysis model to perform bird nest recognition, and a recognition result of the image to be detected is obtained, specifically:

[0034] According to the prompt word text, the image analysis model is used to independently perform a scene description on the relationship between the potential bird's nest target and the substation equipment in each image to be detected, and a description result and a bird's nest target judgment result for each image to be detected are obtained.

[0035] The above solution describes the scene based on the prompt text, taking each image to be detected as a separate scene, outputs the description result and the bird's nest target judgment result, and accurately determines whether there is a bird's nest on the substation equipment.

[0036] In a possible implementation method of the first aspect, outputting the location of the bird's nest in the image to be detected and the corresponding alarm signal according to the recognition result is specifically:

[0037] If the bird's nest target judgment result of the recognition result meets the bird's nest hidden danger characteristics, the alarm signal is sent, and the position, confidence level and scene description of the bird's nest in the image to be detected are displayed through a visual interface.

[0038] When the above solution determines that there is a bird's nest on the current substation equipment, it will issue an alarm signal and visualize the location, confidence level and scenario description of the bird's nest, making it easy to quickly deal with the bird's nest and ensure the safe operation of the substation equipment.

[0039] In a possible implementation method of the first aspect, a drone device is used to collect images in a substation scenario, specifically:

[0040] Set the drone's shooting angle and set the drone's camera to the highest resolution for shooting; wherein the drone's shooting angle includes overhead, side, and overhead shooting;

[0041] Use a drone to photograph the top and sides of substation equipment in a substation scenario, and save the collected images in a lossless compression format.

[0042] This solution uses multiple camera angles to ensure accurate capture of key substation equipment components. To ensure image quality, high-resolution capture is also used, and captured images are saved in a lossless compression format to facilitate subsequent image processing.

[0043] The second aspect of the present application provides a substation bird nest hidden danger identification system based on multi-model collaboration, the system comprising: an image acquisition module, a rapid positioning module, a prompt text generation module, a deep analysis module and a hidden danger alarm module;

[0044] The image acquisition module is used to use drone equipment to collect images in the substation scene to obtain a set of bird nest hidden danger images of the substation;

[0045] The rapid positioning module is used to input the substation bird nest hidden danger image set into a preset target area recognition model to perform potential target recognition and image cropping, and output an image to be detected containing potential bird nest targets;

[0046] The prompt text generation module is used to enhance the visual features of the image to be detected based on a preset recognition task and generate a prompt word text;

[0047] The deep analysis module is used to input the image to be detected and the prompt word text into a preset image analysis model to perform bird nest recognition and obtain a recognition result of the image to be detected;

[0048] The hidden danger warning module is used to output the location of the bird's nest in the image to be detected and the corresponding warning signal according to the recognition result. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solution of the present application, the following is a brief introduction to the drawings required for use in the implementation. Obviously, the drawings described below are only some implementation methods of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0050] Figure 1 This is a schematic diagram of a specific process of a method for identifying bird nest hazards in substations based on multi-model collaboration provided by an embodiment of the present application;

[0051] Figure 2 This is a schematic diagram of high-incidence locations of bird nests in substations, according to a method for identifying bird nest hazards in substations based on multi-model collaboration, provided in one embodiment of the present application;

[0052] Figure 3 This is a target area identification model recognition effect diagram of a substation bird nest hidden danger identification method based on multi-model collaboration provided by a certain embodiment of the present application;

[0053] Figure 4 This is an example diagram of the input and output of a target area identification model for a substation bird nest hazard identification method based on multi-model collaboration provided by an embodiment of the present application;

[0054] Figure 5 This is an example diagram of prompt words for a method for identifying bird nest hazards in substations based on multi-model collaboration provided by a certain embodiment of the present application;

[0055] Figure 6 This is an example image of an input image of an image analysis model for a method for identifying bird nest hazards in substations based on multi-model collaboration, provided in one embodiment of the present application;

[0056] Figure 7 This is a structural diagram of a substation bird nest hazard identification system based on multi-model collaboration provided in a certain embodiment of the present application. DETAILED DESCRIPTION

[0057] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0058] It should be understood that the step numbers used herein are for convenience of description only and are not intended to limit the order in which the steps are performed. In the description of this application, unless otherwise specified, "several" means two or more.

[0059] First embodiment

[0060] Because hay, branches and other materials in bird nests can easily cause line short circuits or grounding faults, leading to equipment damage, power outages and even grid paralysis, accurately identifying bird nests on substation equipment and cleaning them in a timely manner can effectively avoid faults and extend the service life of the equipment. However, existing small model recognition is prone to identifying objects with similar shapes and colors to bird nests as bird nests in complex environments, and large models such as large-scale language models cannot fully utilize the visual information of the original image and lack the ability to extract original image information. Therefore, they cannot accurately determine whether there is a bird nest on the substation equipment in the original image. Based on the above problems, the embodiment of the present application combines a small model with a large model. The small model uses its lightweight characteristics to quickly extract candidate areas for bird nests, and the large model relies on its strong image understanding capabilities to process complex scenes. The two have a clear division of labor, avoiding the shortcomings of a single model in efficiency or accuracy, and fully releasing the technical advantages of large and small models. It can quickly and accurately determine whether there are bird nest hazards in the current substation scene, providing support for maintaining the safe operation of substation equipment.

[0061] like Figure 1 As shown, in order to solve the problem in the prior art that non-bird nest objects with similar colors and shapes cannot be effectively distinguished in complex environments, resulting in a high misidentification rate of substation hidden danger identification, the first embodiment of the present application provides a specific flow chart of a substation bird nest hidden danger identification method based on multi-model collaboration. The substation bird nest hidden danger identification method based on multi-model collaboration of this embodiment includes steps S1 to S4, which are detailed as follows:

[0062] Step S1: Use drone equipment to collect images in a substation scene to obtain a set of substation bird nest hidden danger images.

[0063] In an embodiment of the present application, in order to acquire high-quality images in a complex substation scenario for subsequent identification of bird's nest hazards, multiple drone shooting angles are set, including overhead, side, and overhead shooting, to ensure that key parts such as the top and sides of the substation equipment can be photographed.

[0064] In order to ensure the quality of the collected images, the camera of the drone should be set to the highest resolution, and the multiple collected images of substation bird nest hidden dangers should be stored in a lossless compression format to obtain a substation bird nest hidden danger image set.

[0065] Figure 2 These images, taken by drone, show high-incidence locations for bird nests in substations. Four common locations for bird nests are marked with red rectangles: at the connection points of substation switch blades, near the transformer oil level gauge, in transformer grooves, and on the structure. These images demonstrate that bird nests in substations are small, occur in a variety of locations, and have a wide variety of shapes.

[0066] Step S2: input the substation bird nest hidden danger image set into a preset target area recognition model to perform potential target recognition and image cropping, and output an image to be detected containing potential bird nest targets.

[0067] Considering that the target area recognition model is difficult to accurately capture the core features of bird nests in complex scenes, it is unable to effectively distinguish non-bird nest objects with similar colors and shapes such as rust, and it is difficult to identify interference elements such as the grass background in the hollow part of the equipment. Therefore, the embodiment of the present application first uses the target area recognition model to quickly identify potential bird nest targets suspected of bird nests in the substation bird nest hazard image set, and cuts out these potential bird nest targets to provide data support for subsequent in-depth analysis.

[0068] Since the target area recognition model is used to preliminarily identify areas where bird nest targets may exist, the model is lightweight and belongs to a small model, so the small model will be used to refer to the target area recognition model in the future; the subsequent image analysis model has a more complex structure and belongs to a large-scale language model, so it will be referred to as a large model in the future.

[0069] Figure 3 The following diagram shows the effect of using only small models to identify bird nests in a complex substation scenario. Figure 3 In the first two pictures on the left, the small model easily misidentifies the rust and other objects with similar colors and shapes as bird nests; Figure 3 In the last two pictures on the right, the small model easily misidentifies the grass background at the hollowed-out position of the substation equipment as a bird's nest.

[0070] Commonly used small models are generally YOLO series network architectures. Optionally, the embodiment of the present application uses a small model based on YOLOv12 to quickly locate potential bird nest targets. In other embodiments, small models based on YOLOv8 and YOLOv10 can also be used.

[0071] In order to improve the potential bird's nest target recognition effect of the small model, before the substation bird's nest hidden danger image set is input into the preset small model, the substation bird's nest hidden danger image set is also preprocessed to improve the image quality and improve the recognition effect of the small model.

[0072] Specifically, the system sequentially performs image enhancement, horizontal flipping, image transformation, visual adjustment, and object cloning on a set of substation bird nest hazard images. Image enhancement involves stitching four images into a single image to simulate complex scenes, improving the small model's multi-scale object detection capabilities. Image transformation involves performing image processing such as translation and scaling. Visual adjustment primarily adjusts the image's hue, saturation, and brightness. Object cloning involves copying and pasting objects from other images into the current image, addressing the long-tail distribution of bird nest morphology.

[0073] Exemplarily, for visual adjustment, an embodiment of the present application performs a ±1.5% hue shift on the substation bird nest hazard image set to simulate changes in light color temperature; sets the saturation scaling coefficient to 0.7 for saturation adjustment to enhance the robustness of low-saturation scenes; and sets the brightness scaling coefficient to 0.4 to cope with overexposed or underexposed environments.

[0074] The long-tail distribution problem refers to data distributions in which a few categories or events occur very frequently, while the majority have a very low probability of occurrence. This distributional characteristic causes the head (high-frequency portion) to receive the majority of attention and resources, while the tail (low-frequency portion) contains a large number of categories that individually occur less frequently but are large in total. This long-tail distribution problem leads to an uneven data distribution, causing the model to favor high-frequency categories and ignore low-frequency categories during training. Furthermore, the sparseness of low-frequency data makes it difficult for the model to learn and generalize, especially when resources are limited.

[0075] In order to cope with the above situation, the embodiment of the present application performs target replication on the substation bird nest hidden danger image set, mainly copying the bird nest target to other images, increasing the proportion of images containing bird nest targets, solving the problem of uneven data distribution, and balancing the samples of high-frequency and low-frequency categories.

[0076] The small model based on YOLOv12 used in the embodiment of the present application is mainly composed of a backbone network, a neck network and a head network.

[0077] The backbone network is responsible for basic feature extraction. Its core consists of convolutional layers, R-ELAN modules, and A2 area attention modules:

[0078] (1) Convolutional layer: uses standard convolution combined with BN (batch normalization) and SiLU activation function to extract low-level image features (such as edges and textures) and generate corresponding feature maps;

[0079] (2) R-ELAN module: Based on the improvement of ELAN, it introduces residual connection and layer scaling technology to optimize the gradient flow of deep networks. It aggregates features of different receptive fields through a multi-branch structure to improve feature expression capabilities. ELAN is the full name of Efficient Long-Distance Attention Network, also known as Efficient Long-Distance Attention Network.

[0080] (3) A2 Regional Attention Module: This module divides the feature map into local strips by height or width, independently calculates attention weights, and dynamically enhances the feature response of key areas. Compared to global attention, this module reduces computational complexity by approximately 50%, while retaining a large receptive field and significantly reducing memory usage.

[0081] The neck network is used for multi-scale feature fusion and mainly includes a splicing layer, a learnable upsampling module, and a C2f module:

[0082] (1) Concatenation layer: It integrates features from different layers of the backbone network, combining shallow details (such as small object outlines) with deep semantic information (such as object categories);

[0083] (2) Learnable upsampling module: Transposed convolution is used instead of traditional interpolation to accurately restore the feature map resolution and improve the detection of small targets. After upsampling, the channel dimension is adjusted through lightweight convolution to reduce redundancy.

[0084] (3) C2f module: Cross-stage feature fusion design, which divides the input features into the main branch and the residual branch. The main branch retains the original features, and the residual branch stacks lightweight convolutional layers, and the end splices and fuses multi-scale information.

[0085] The head network adopts a decoupled design, including a separate classification module and a positioning task module:

[0086] (1) Separation classification module: outputs the target category probability and uses Focal Loss to alleviate the category imbalance problem;

[0087] (2) Positioning task module: predict bounding box coordinates and confidence, introduce DFL (distributed focus loss) to improve positioning accuracy; accurately allocate positive and negative samples based on dynamic allocation strategy.

[0088] The small model was trained using drone-collected images of bird nest hazards at substations as input. During training, each batch of eight images was trained on a single card, pre-loaded with parameters trained on MS-COCO object detection. The total number of training cycles was 100, and the input image size was 1920×1920.

[0089] The pre-processed substation bird nest hidden danger image set is input into the trained small model to identify potential targets suspected of bird nests, and each potential bird nest target is marked with a rectangular box.

[0090] For example, let the coordinates of a potential bird's nest rectangle be (x min ,y min ,x max ,y max ), calculate the geometric center point of the rectangle (x center ,y center ). Then calculate the width and height of the rectangular frame and enlarge the width and height of the rectangular frame according to the set ratio. In the embodiment of the present application, the ratio is set to 5 times. Take the maximum value of the enlarged width and height as the side length, and use the geometric center point (x center ,y center ) as the center to build a square frame and get the square frame coordinates (x1 min ,y1 min ,x1 max ,y1 max ). Then perform boundary check on the square frame coordinates and detect x1 min and y1 min Is it greater than or equal to 0, and limit the square frame to not exceed the image size, that is, x1 max To be less than or equal to the image width and y1 max Finally, only the content within the square frame is retained, and the substation bird nest hidden danger image is cropped to obtain the corresponding image to be detected.

[0091] Figure 4 The following diagrams show the input and output of the small model: (a) shows the input image of the small model, and (b) shows the output image of the small model. In (b), a red rectangle is used to mark the identified potential bird's nest. This rectangle can be used to crop (b) to obtain the corresponding image to be detected.

[0092] Step S3: Based on the preset recognition task, visual features of the image to be detected are enhanced to generate prompt word text.

[0093] In order to improve the image understanding ability of the large model in complex scenes, the embodiment of the present application provides prompt word text as the basis for image recognition of the large model, so that the large model can fully utilize the information of the prompt word text and accurately analyze the image.

[0094] In the embodiment of the present application, the prompt word text is constructed for the large model mainly based on the role positioning and recognition tasks, the construction of the judgment standard system, the guidance of visual feature enhancement, the analysis process specification and the output format constraint, specifically:

[0095] (1) Role positioning and identification task: The role of the extracted prompt word is set as "experienced hidden danger inspector" to strengthen the large model's professional understanding of the substation. Then set the identification task to help the large model focus on the identification of specific objects to distinguish them from other defects and highlight the specificity of the detection target. For example, if the identification task is set to identify hay, then the large model will focus on identifying whether there is hay in the image and the spatial relationship between hay and substation equipment through features such as the color of the hay; in addition, the identification task can also be set to identify rust;

[0096] (2) Construction of judgment standard system: The identification of potential bird nest targets in the image is mainly based on the existence and position relationship. For example, if the recognition task is to identify hay, then it will be determined whether there is hay in the image and what the spatial relationship between hay and substation equipment is (for example, hay is located on substation equipment), and the corresponding judgment matrix will be constructed accordingly; in addition, a confidence control is set. When the existence or spatial relationship of the potential bird nest target cannot be determined, the "unable to determine" option can be selected to avoid the large model forcibly outputting incorrect results when the data quality is insufficient;

[0097] (3) Visual feature enhancement guidance: Here, the visual features of potential bird nest targets in the image to be detected are enhanced, including color features, material features, and morphological features. Taking hay as an example, the yellow characteristics of hay are clearly defined, forming a sharp contrast with the reddish-brown color of rust; the metal surface properties of substation equipment are emphasized, and "walls do not belong to equipment" is emphasized to help the large model establish the identification benchmark for equipment and non-equipment areas; morphological interference items are eliminated by "hay does not include brooms" and "rust is not hay" to emphasize morphological features;

[0098] (4) Analysis process specifications: The large model is required to first scan the scene as a whole and then focus on the device area, which is in line with the cognitive laws of manual visual inspection; the large model is forced to describe first and then judge, ensuring that the judgment conclusion is supported by clear visual evidence; finally, the statement of "each image is analyzed independently" is used to avoid the model from generating context-dependent bias, that is, each time the large model describes and analyzes the image to be detected, it is described and analyzed independently through the corresponding prompt word text, which is unrelated to the previous description and analysis;

[0099] (5) Output format constraints: A dual-field format of “description and judgment” is used to achieve standardized output of detection results; the requirement of “must be completely consistent” is used to ensure that the judgment results accurately match the preset classifications; the description field provides a traceable visual basis for the judgment conclusions, enhancing the credibility of the results, that is, the description results of the image to be detected must provide a visual basis for the judgment results of the bird’s nest target.

[0100] For example, taking the task of identifying hay as an example, Figure 5The resulting prompt text based on the above rules is shown. The judgment criteria outline the existence and positional relationships of four target types, requiring exact consistency. Otherwise, the "Cannot Determine" option is selected. The large model is then prompted to use visual features to determine whether the identified target matches the characteristics of hay and whether any substation equipment has been misidentified. Finally, the large model is instructed to output a scenario description and the corresponding judgment results.

[0101] Based on the above prompt text, the large model can describe whether there is hay in the image, the location of the hay, etc., and if hay exists, whether the hay is on the substation equipment, providing evidence for issuing a bird's nest hidden danger alarm signal.

[0102] Step S4: input the image to be detected and the prompt word text into a preset image analysis model to perform bird nest recognition, and obtain a recognition result of the image to be detected.

[0103] In an embodiment of the present application, a large model based on Qwen2-VL-7B is used to perform in-depth analysis of the image to be detected, relying on the image understanding ability of the large model to process feature recognition in complex scenes, avoiding the shortcomings of a single model in efficiency or accuracy, and fully releasing the technical advantages of large and small models. Qwen2-VL-7B is a type of Qwen2-VL large model. The Qwen2-VL large model is a multimodal large language model that can process text, images, multiple images and video inputs, and is particularly good at visual-language tasks. The core architecture of the Qwen2-VL large model adopts the architectural design of dynamic resolution visual encoder and multimodal rotation position encoding. It realizes efficient parsing of images of any size by dynamically adjusting the number of visual tokens. Combined with multimodal rotation position encoding technology, it synchronously models image spatial features and video timing information, significantly improving recognition accuracy in complex scenes.

[0104] The large model commonly used in the prior art is the CLIP large model, which achieves coarse-grained matching of images and text through comparative learning, but has limited semantic parsing capabilities for complex scenes. As an improvement to this solution, the embodiment of the present application specifically optimizes the Qwen2-VL large model for the identification scenario of bird nest hazards in substations. By constructing a multimodal dataset containing a large number of substation equipment images and performing QLoRA fine-tuning technology, the model can better adapt to the complex environment of the substation. Compared with the commonly used CLIP large model, the Qwen2-VL large model can accurately parse information such as semantics and depth of images, reducing misidentification of rust, hollowed-out grass locations, and other misjudgments.

[0105] QLoRA fine-tuning technology freezes the backbone network parameters of the large model and trains only the low-rank adapter. While maintaining model versatility, it optimizes parameters for the bird's nest scenario from the perspective of a substation drone. QLoRA fine-tuning also improves model performance with limited computing resources through dynamic learning rate scheduling and gradient accumulation.

[0106] The acquired images to be detected containing potential bird's nest targets and the prompt word text corresponding to each image are input into the trained large model based on Qwen2-VL-7B. Through the dynamic resolution visual encoder and multimodal rotation position encoding architecture of the large model, the image visual information is fully utilized to independently describe the relationship between the potential bird's nest targets and substation equipment in each image to be detected. Finally, the description results of each image to be detected and the bird's nest target judgment results are output.

[0107] For example, Figure 6 The image to be detected is input into the large model in a certain image recognition in the embodiment of the present application, including Figure (a), Figure (b) and Figure (c). The following Table 1 is an example of the input and output of the large model.

[0108] Table 1 Image input and text output of the large model

[0109]

[0110]

[0111] The collaborative cascade structure of small and large models achieves a combination of rapid positioning and in-depth understanding. The small model performs preliminary target detection and quickly screens possible target areas. The large model then deeply understands and analyzes the output of the small model to accurately determine whether bird nest hazards are located in substation equipment. This allows for better precision identification in complex substation environments, reducing misjudgments and missed detections.

[0112] Step S5: outputting the location of the bird's nest in the image to be detected and the corresponding alarm signal according to the recognition result.

[0113] The system determines whether to issue an alarm based on the description output by the large model and the bird's nest target identification result. For example, if the large model output of a certain image to be detected indicates the presence of hay on the equipment, it indicates the presence of a bird's nest on the substation equipment, and an alarm signal and the corresponding nest location should be issued; otherwise, no alarm is issued.

[0114] As an improvement to the above solution, the embodiment of the present application also provides a visual interface to display the location, confidence level and scene description of the bird's nest, to assist maintenance personnel in quickly locating the bird's nest and performing maintenance in a timely manner.

[0115] The implementation of the embodiments of the present application has the following beneficial effects:

[0116] In this embodiment, a drone is used to photograph substation equipment from multiple viewing angles in a substation scenario, generating a set of substation bird nest hazard images. A small model is first used to locate potential bird nests in the captured images. Taking advantage of the small model's lightweight nature, areas suspected of being bird nests are quickly extracted. Areas irrelevant to nest detection are then cropped to eliminate interference, resulting in an image containing potential bird nests. Prompt text is then extracted from the image to be detected to assist the large model in image recognition, enabling the large model to fully utilize the visual features of potential bird nests and better adapt to the complex substation environment. A large model capable of handling complex scenarios is used to perform in-depth analysis of the image to be detected, distinguishing non-bird nest objects with similar color and shape within the image based on their visual features. This allows for filtering out misidentifications and accurately determining the presence of bird nests on substation equipment. Each image to be detected is treated as a separate scene and outputs a corresponding recognition result. Finally, based on the obtained recognition results, the bird nest is accurately located and an alarm signal is sent, significantly improving the accuracy of substation hazard identification and the safety of substation operations.

[0117] Second embodiment

[0118] Furthermore, in order to implement the substation bird nest hidden danger identification system based on multi-model collaboration corresponding to the above method embodiment to achieve the corresponding functions and technical effects, Figure 7 A structural diagram of a substation bird nest hazard identification system based on multi-model collaboration is provided. For ease of explanation, only the parts related to this embodiment are shown. The substation bird nest hazard identification system based on multi-model collaboration provided by the embodiment of the present application includes:

[0119] The image acquisition module 201 is used to use drone equipment to collect images in the substation scene to obtain a set of images of bird nest hazards in the substation.

[0120] In an embodiment of the present application, in order to acquire high-quality images in a complex substation scenario for subsequent identification of bird's nest hazards, multiple drone shooting angles are set, including overhead, side, and overhead shooting, to ensure that key parts such as the top and sides of the substation equipment can be photographed.

[0121] In order to ensure the quality of the collected images, the camera of the drone should be set to the highest resolution, and the multiple collected images of substation bird nest hidden dangers should be stored in a lossless compression format to obtain a substation bird nest hidden danger image set.

[0122] The rapid positioning module 202 is used to input the substation bird nest hidden danger image set into a preset target area recognition model to perform potential target recognition and image cropping, and output an image to be detected containing potential bird nest targets.

[0123] In the embodiment of the present application, potential bird nest targets are identified from the substation bird nest hidden danger image set by the target area recognition model, and each potential bird nest target is marked with a rectangular frame;

[0124] Enlarging the rectangular frame at a preset ratio, and taking the maximum value of the width and height of the enlarged rectangular frame as the side length to generate a corresponding square frame;

[0125] With the goal of retaining only the content within the square frame, the substation bird nest hidden danger image set is cropped to output the image to be detected.

[0126] The prompt text generation module 203 is used to enhance the visual features of the image to be detected based on a preset recognition task and generate a prompt word text.

[0127] In an embodiment of the present application, based on a preset recognition task, the dimension judgment and visual feature enhancement of the potential bird's nest target in the image to be detected are performed, and a corresponding prompt text is generated for each image to be detected; wherein, the recognition task includes identifying hay and identifying rust.

[0128] The deep analysis module 204 is used to input the image to be detected and the prompt word text into a preset image analysis model to perform bird nest recognition and obtain a recognition result of the image to be detected.

[0129] In an embodiment of the present application, a large model based on Qwen2-VL-7B is used to perform in-depth analysis of the image to be detected, relying on the image understanding ability of the large model to process feature recognition in complex scenes, avoiding the shortcomings of a single model in efficiency or accuracy, and fully releasing the technical advantages of large and small models. Qwen2-VL-7B is a type of Qwen2-VL large model. The Qwen2-VL large model is a multimodal large language model that can process text, images, multiple images and video inputs, and is particularly good at visual-language tasks. The core architecture of the Qwen2-VL large model adopts the architectural design of dynamic resolution visual encoder and multimodal rotation position encoding. It realizes efficient parsing of images of any size by dynamically adjusting the number of visual tokens. Combined with multimodal rotation position encoding technology, it synchronously models image spatial features and video timing information, significantly improving recognition accuracy in complex scenes.

[0130] The large model commonly used in the prior art is the CLIP large model, which achieves coarse-grained matching of images and text through comparative learning, but has limited semantic parsing capabilities for complex scenes. As an improvement to this solution, the embodiment of the present application specifically optimizes the Qwen2-VL large model for the identification scenario of bird nest hazards in substations. By constructing a multimodal dataset containing a large number of substation equipment images and performing QLoRA fine-tuning technology, the model can better adapt to the complex environment of the substation. Compared with the commonly used CLIP large model, the Qwen2-VL large model can accurately parse information such as semantics and depth of images, reducing misidentification of rust, hollowed-out grass locations, and other misjudgments.

[0131] QLoRA fine-tuning technology freezes the backbone network parameters of the large model and trains only the low-rank adapter. While maintaining model versatility, it optimizes parameters for the bird's nest scenario from the perspective of a substation drone. QLoRA fine-tuning also improves model performance with limited computing resources through dynamic learning rate scheduling and gradient accumulation.

[0132] The acquired images to be detected containing potential bird's nest targets and the prompt word text corresponding to each image are input into the trained large model based on Qwen2-VL-7B. Through the dynamic resolution visual encoder and multimodal rotation position encoding architecture of the large model, the image visual information is fully utilized to independently describe the relationship between the potential bird's nest targets and substation equipment in each image to be detected. Finally, the description results of each image to be detected and the bird's nest target judgment results are output.

[0133] The collaborative cascade structure of small and large models achieves a combination of rapid positioning and in-depth understanding. The small model performs preliminary target detection and quickly screens possible target areas. The large model then deeply understands and analyzes the output of the small model to accurately determine whether bird nest hazards are located in substation equipment. This allows for better precision identification in complex substation environments, reducing misjudgments and missed detections.

[0134] The hidden danger warning module 205 is used to output the location of the bird's nest in the image to be detected and the corresponding warning signal according to the recognition result.

[0135] The system determines whether to issue an alarm based on the description output by the large model and the bird's nest target identification result. For example, if the large model output of a certain image to be detected indicates the presence of hay on the equipment, it indicates the presence of a bird's nest on the substation equipment, and an alarm signal and the corresponding nest location should be issued; otherwise, no alarm is issued.

[0136] As an improvement to the above solution, the embodiment of the present application also provides a visual interface to display the location, confidence level and scene description of the bird's nest, to assist maintenance personnel in quickly locating the bird's nest and performing maintenance in a timely manner.

[0137] In some embodiments, the rapid positioning module 202 is specifically:

[0138] Considering that the target area recognition model is difficult to accurately capture the core features of bird nests in complex scenes, it is unable to effectively distinguish non-bird nest objects with similar colors and shapes such as rust, and it is difficult to identify interference elements such as the grass background in the hollow part of the equipment. Therefore, the embodiment of the present application first uses the target area recognition model to quickly identify potential bird nest targets suspected of bird nests in the substation bird nest hidden danger image set, and cuts out these potential bird nest targets to provide data support for subsequent in-depth analysis.

[0139] Since the target area recognition model is used to preliminarily identify areas where bird nest targets may exist, the model is lightweight and belongs to a small model, so the small model will be used to refer to the target area recognition model in the future; the subsequent image analysis model has a more complex structure and belongs to a large-scale language model, so it will be referred to as a large model in the future.

[0140] Commonly used small models are generally YOLO series network architectures. Optionally, the embodiment of the present application uses a small model based on YOLOv12 to quickly locate potential bird nest targets. In other embodiments, small models based on YOLOv8 and YOLOv10 can also be used.

[0141] In order to improve the potential bird's nest target recognition effect of the small model, before the substation bird's nest hidden danger image set is input into the preset small model, the substation bird's nest hidden danger image set is also preprocessed to improve the image quality and improve the recognition effect of the small model.

[0142] Specifically, the system sequentially performs image enhancement, horizontal flipping, image transformation, visual adjustment, and object cloning on a set of substation bird nest hazard images. Image enhancement involves stitching four images into a single image to simulate complex scenes, improving the small model's multi-scale object detection capabilities. Image transformation involves performing image processing such as translation and scaling. Visual adjustment primarily adjusts the image's hue, saturation, and brightness. Object cloning involves copying and pasting objects from other images into the current image, addressing the long-tail distribution of bird nest morphology.

[0143] Exemplarily, for visual adjustment, an embodiment of the present application performs a ±1.5% hue shift on the substation bird nest hazard image set to simulate changes in light color temperature; sets the saturation scaling coefficient to 0.7 for saturation adjustment to enhance the robustness of low-saturation scenes; and sets the brightness scaling coefficient to 0.4 to cope with overexposed or underexposed environments.

[0144] The long-tail distribution problem refers to data distributions in which a few categories or events occur very frequently, while the majority have a very low probability of occurrence. This distributional characteristic causes the head (high-frequency portion) to receive the majority of attention and resources, while the tail (low-frequency portion) contains a large number of categories that individually occur less frequently but are large in total. This long-tail distribution problem leads to an uneven data distribution, causing the model to favor high-frequency categories and ignore low-frequency categories during training. Furthermore, the sparseness of low-frequency data makes it difficult for the model to learn and generalize, especially when resources are limited.

[0145] In order to cope with the above situation, the embodiment of the present application performs target replication on the substation bird nest hidden danger image set, mainly copying the bird nest target to other images, increasing the proportion of images containing bird nest targets, solving the problem of uneven data distribution, and balancing the samples of high-frequency and low-frequency categories.

[0146] The small model based on YOLOv12 used in the embodiment of the present application is mainly composed of a backbone network, a neck network and a head network.

[0147] The backbone network is responsible for basic feature extraction. Its core consists of convolutional layers, R-ELAN modules, and A2 area attention modules:

[0148] (4) Convolutional layer: uses standard convolution combined with BN (batch normalization) and SiLU activation function to extract low-level features of the image (such as edges and textures) and generate corresponding feature maps;

[0149] (5) R-ELAN module: Based on the improvement of ELAN, it introduces residual connection and layer scaling technology to optimize the gradient flow of deep networks. It aggregates features of different receptive fields through a multi-branch structure to improve feature expression capabilities. ELAN is the full name of Efficient Long-Distance Attention Network, also known as Efficient Long-Distance Attention Network.

[0150] (6) A2 Regional Attention Module: This module divides the feature map into local strips by height or width, independently calculates attention weights, and dynamically enhances the feature response of key areas. Compared to global attention, the computational complexity is reduced by approximately 50%, while retaining a large receptive field and significantly reducing memory usage.

[0151] The neck network is used for multi-scale feature fusion and mainly includes a splicing layer, a learnable upsampling module, and a C2f module:

[0152] (4) Splicing layer: It integrates features from different levels of the backbone network, combining shallow details (such as small object outlines) with deep semantic information (such as object categories);

[0153] (5) Learnable upsampling module: Transposed convolution is used instead of traditional interpolation to accurately restore the feature map resolution and improve the detection of small targets. After upsampling, the channel dimension is adjusted through lightweight convolution to reduce redundancy.

[0154] (6) C2f module: Cross-stage feature fusion design, which divides the input features into the main branch and the residual branch. The main branch retains the original features, and the residual branch stacks lightweight convolutional layers, and the end splices and fuses multi-scale information.

[0155] The head network adopts a decoupled design, including a separate classification module and a positioning task module:

[0156] (3) Separation classification module: outputs the target category probability and uses Focal Loss to alleviate the category imbalance problem;

[0157] (4) Positioning task module: predict bounding box coordinates and confidence, introduce DFL (distributed focus loss) to improve positioning accuracy; accurately allocate positive and negative samples based on dynamic allocation strategy.

[0158] The small model was trained using drone-collected images of bird nest hazards at substations as input. During training, each batch of eight images was trained on a single card, pre-loaded with parameters trained on MS-COCO object detection. The total number of training cycles was 100, and the input image size was 1920×1920.

[0159] The pre-processed substation bird nest hidden danger image set is input into the trained small model to identify potential targets suspected of bird nests, and each potential bird nest target is marked with a rectangular box.

[0160] For example, let the coordinates of a potential bird's nest rectangle be (x min ,y min ,x max ,y max ), calculate the geometric center point of the rectangle (x center ,y center ). Then calculate the width and height of the rectangular frame and enlarge the width and height of the rectangular frame according to the set ratio. In the embodiment of the present application, the ratio is set to 5 times. Take the maximum value of the enlarged width and height as the side length, and use the geometric center point (x center ,y center ) as the center to build a square frame and get the square frame coordinates (x1 min ,y1 min ,x1 max ,y1 max ). Then perform boundary check on the square frame coordinates and detect x1 min and y1 min Is it greater than or equal to 0, and limit the square frame to not exceed the image size, that is, x1max To be less than or equal to the image width and y1 max Finally, only the content within the square frame is retained, and the substation bird nest hidden danger image is cropped to obtain the corresponding image to be detected.

[0161] In some embodiments, the prompt text generation module 203 is specifically:

[0162] In order to improve the image understanding ability of the large model in complex scenes, the embodiment of the present application provides prompt word text as the basis for image recognition of the large model, so that the large model can fully utilize the information of the prompt word text and accurately analyze the image.

[0163] In the embodiment of the present application, the prompt word text is constructed for the large model mainly based on the role positioning and recognition tasks, the construction of the judgment standard system, the guidance of visual feature enhancement, the analysis process specification and the output format constraint, specifically:

[0164] (1) Role positioning and identification task: The role of the extracted prompt word is set as "experienced hidden danger inspector" to strengthen the large model's professional understanding of the substation. Then set the identification task to help the large model focus on the identification of specific objects to distinguish them from other defects and highlight the specificity of the detection target. For example, if the identification task is set to identify hay, then the large model will focus on identifying whether there is hay in the image and the spatial relationship between hay and substation equipment through features such as the color of the hay; in addition, the identification task can also be set to identify rust;

[0165] (2) Construction of judgment standard system: The identification of potential bird nest targets in the image is mainly based on the existence and position relationship. For example, if the recognition task is to identify hay, then it will be determined whether there is hay in the image and what the spatial relationship between hay and substation equipment is (for example, hay is located on substation equipment), and the corresponding judgment matrix will be constructed accordingly; in addition, a confidence control is set. When the existence or spatial relationship of the potential bird nest target cannot be determined, the "unable to determine" option can be selected to avoid the large model forcibly outputting incorrect results when the data quality is insufficient;

[0166] (3) Visual feature enhancement guidance: Here, the visual features of potential bird nest targets in the image to be detected are enhanced, including color features, material features, and morphological features. Taking hay as an example, the yellow characteristics of hay are clearly defined, forming a sharp contrast with the reddish-brown color of rust; the metal surface properties of substation equipment are emphasized, and "walls do not belong to equipment" is emphasized to help the large model establish the identification benchmark for equipment and non-equipment areas; morphological interference items are eliminated by "hay does not include brooms" and "rust is not hay" to emphasize morphological features;

[0167] (4) Analysis process specifications: The large model is required to first scan the scene as a whole and then focus on the device area, which is in line with the cognitive laws of manual visual inspection; the large model is forced to describe first and then judge, ensuring that the judgment conclusion is supported by clear visual evidence; finally, the statement of "each image is analyzed independently" is used to avoid the model from generating context-dependent bias, that is, each time the large model describes and analyzes the image to be detected, it is described and analyzed independently through the corresponding prompt word text, which is unrelated to the previous description and analysis;

[0168] (5) Output format constraints: A dual-field format of “description and judgment” is used to achieve standardized output of detection results; the requirement of “must be completely consistent” is used to ensure that the judgment results accurately match the preset classifications; the description field provides a traceable visual basis for the judgment conclusions, enhancing the credibility of the results, that is, the description results of the image to be detected must provide a visual basis for the judgment results of the bird’s nest target.

[0169] Based on the above prompt text, the large model can describe whether there is hay in the image, the location of the hay, etc., and if hay exists, whether the hay is on the substation equipment, providing evidence for issuing a bird's nest hidden danger alarm signal.

[0170] The implementation of the embodiments of the present application has the following beneficial effects:

[0171] In this embodiment, a drone is used to photograph substation equipment from multiple viewing angles in a substation scenario, generating a set of substation bird nest hazard images. A small model is first used to locate potential bird nests in the captured images. Taking advantage of the small model's lightweight nature, areas suspected of being bird nests are quickly extracted. Areas irrelevant to nest detection are then cropped to eliminate interference, resulting in an image containing potential bird nests. Prompt text is then extracted from the image to be detected to assist the large model in image recognition, enabling the large model to fully utilize the visual features of potential bird nests and better adapt to the complex substation environment. A large model capable of handling complex scenarios is used to perform in-depth analysis of the image to be detected, distinguishing non-bird nest objects with similar color and shape within the image based on their visual features. This allows for filtering out misidentifications and accurately determining the presence of bird nests on substation equipment. Each image to be detected is treated as a separate scene and outputs a corresponding recognition result. Finally, based on the obtained recognition results, the bird nest is accurately located and an alarm signal is sent, significantly improving the accuracy of substation hazard identification and the safety of substation operations.

[0172] The specific embodiments described above further illustrate the purpose, technical solutions, and beneficial effects of this application. It should be understood that the above description is merely a specific embodiment of this application and is not intended to limit the scope of protection of this application. In particular, it should be noted that for those skilled in the art, any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of this application should be included in the scope of protection of this application.

Claims

1. A method for identifying bird nest hazards in substations based on multi-model collaboration, characterized in that: include: Use drone equipment to collect images in substation scenes and obtain a set of images of bird nest hazards in substations; Input the substation bird nest hidden danger image set into a preset target area recognition model to perform potential target recognition and image cropping, and output an image to be detected containing potential bird nest targets; Based on a preset recognition task, visual features of the image to be detected are enhanced to generate prompt word text; Inputting the image to be detected and the prompt word text into a preset image analysis model to perform bird nest recognition, and obtaining a recognition result of the image to be detected; According to the recognition result, the location of the bird's nest in the image to be detected and the corresponding alarm signal are output.

2. The method for identifying bird nest hazards in substations based on multi-model collaboration according to claim 1 is characterized in that: The substation bird nest hidden danger image set is input into a preset target area recognition model to perform potential target recognition and image cropping, and an image to be detected containing potential bird nest targets is output, specifically: Identifying potential bird nest targets from the substation bird nest hidden danger image set using the target area recognition model, and marking each potential bird nest target with a rectangular frame; Enlarging the rectangular frame at a preset ratio, and taking the maximum value of the width and height of the enlarged rectangular frame as the side length to generate a corresponding square frame; With the goal of retaining only the content within the square frame, the substation bird nest hidden danger image set is cropped to output the image to be detected.

3. The method for identifying bird nest hazards in substations based on multi-model collaboration according to claim 2 is characterized in that: The identifying of potential bird nest targets from the substation bird nest hidden danger image set by the target area recognition model is specifically as follows: Extracting low-level image features from the substation bird nest hidden danger image set and generating a corresponding feature map; wherein the low-level image features include edges and textures of objects in the image; By performing feature splicing, upsampling and feature fusion on the feature map, potential bird nest targets in the substation bird nest hidden danger image set are identified.

4. The method for identifying bird nest hazards in substations based on multi-model collaboration according to claim 2 is characterized in that: Before inputting the substation bird nest hidden danger image set into a preset target area recognition model, the method further includes: Performing image enhancement, horizontal flipping, image transformation, visual adjustment, and target copying on the substation bird nest hidden danger image set in sequence; Among them, the image enhancement is to stitch several images into one image; the image transformation is to translate and scale the image; the visual adjustment is to adjust the hue, saturation and brightness of the image; the target copying is to copy and paste the target in other images to the current image.

5. The method for identifying bird nest hazards in substations based on multi-model collaboration according to claim 1 is characterized in that: Based on the preset recognition task, the visual features of the image to be detected are enhanced to generate prompt word text, specifically: Based on a preset recognition task, the dimension judgment and visual feature enhancement of the potential bird's nest target in the image to be detected are performed, and corresponding prompt text is generated for each image to be detected; wherein, the recognition task includes identifying hay and identifying rust.

6. The method for identifying bird nest hazards in substations based on multi-model collaboration according to claim 5 is characterized in that: Based on the preset recognition task, the dimension judgment and visual feature enhancement of the potential bird's nest target in the image to be detected are performed, and corresponding prompt text is generated for each image to be detected, specifically: Constructing a judgment matrix based on the spatial relationship between the identification target of the identification task and the substation equipment, and the existence of the identification target; According to the recognition task, the color features, material features and morphological features of the potential bird's nest target in the image to be detected are enhanced; The image to be detected after feature enhancement is combined with the judgment matrix to generate the prompt text.

7. The method for identifying bird nest hazards in substations based on multi-model collaboration according to claim 1 is characterized in that: The image to be detected and the prompt word text are input into a preset image analysis model to perform bird nest recognition, and a recognition result of the image to be detected is obtained, specifically: According to the prompt word text, the image analysis model is used to independently perform a scene description on the relationship between the potential bird's nest target and the substation equipment in each image to be detected, and a description result and a bird's nest target judgment result for each image to be detected are obtained.

8. The method for identifying bird nest hazards in substations based on multi-model collaboration according to claim 1 is characterized in that: Outputting the bird's nest location in the image to be detected and the corresponding alarm signal according to the recognition result is specifically: If the bird's nest target judgment result of the recognition result meets the bird's nest hidden danger characteristics, the alarm signal is sent, and the position, confidence level and scene description of the bird's nest in the image to be detected are displayed through a visual interface.

9. The method for identifying bird nest hazards in substations based on multi-model collaboration according to claim 1 is characterized in that: The use of drone equipment to collect images in a substation scene is specifically as follows: Set the drone's shooting angle and set the drone's camera to the highest resolution for shooting; wherein the drone's shooting angle includes overhead, side, and overhead shooting; The top and sides of substation equipment are photographed by drone in a substation scene, and the collected images are saved in a lossless compression format.

10. A substation bird nest hidden danger identification system based on multi-model collaboration, characterized by: include: Image acquisition module, rapid positioning module, prompt text generation module, deep analysis module and hidden danger warning module; The image acquisition module is used to use drone equipment to collect images in the substation scene to obtain a set of bird nest hidden danger images of the substation; The rapid positioning module is used to input the substation bird nest hidden danger image set into a preset target area recognition model to perform potential target recognition and image cropping, and output an image to be detected containing potential bird nest targets; The prompt text generation module is used to enhance the visual features of the image to be detected based on a preset recognition task and generate a prompt word text; The deep analysis module is used to input the image to be detected and the prompt word text into a preset image analysis model to perform bird nest recognition and obtain a recognition result of the image to be detected; The hidden danger warning module is used to output the location of the bird's nest in the image to be detected and the corresponding warning signal according to the recognition result.

Citation Information

Patent Citations

  • Behavior recognition method and device based on multi-modal large model and electronic equipment

    CN118314624A

  • Zero-sample potential risk behavior detection method and device based on multi-modal large model

    CN118570868A

  • Iron tower bird nest monitoring method and system based on multi-modal large model

    CN118692028A

  • YOLO and CLIP-based bird nest identification method and device

    CN118898779A

Cited By

  • Hidden danger early warning method, device and equipment for construction site and medium

    CN121074810A

  • Traffic illegal behavior identification method and device

    CN121214367A