Insulator defect detection method and system based on multi-modal large model tool invocation

CN122821158APending Publication Date: 2026-09-25CHENGDU UNIV OF INFORMATION TECH +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610996815.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-06
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0008]本发明的目的是提供一种基于多模态大模型工具调用的绝缘子缺陷检测方法及系统,旨在解决现有技术中变电站绝缘子小尺度、弱纹理缺陷检测精度低、传统检测模型无法主动聚焦局部细节、通用多模态大模型直接应用检测效果不稳定的问题

Benefits of technology

1、构建了适配多轮推理的高质量数据集并实现高效领域适配

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821158A_ABST
    Figure CN122821158A_ABST
Patent Text Reader

Abstract

The application discloses an insulator defect detection method and system based on multi-modal large model tool calling, and relates to the technical field of insulator defect detection. The application firstly constructs a multi-modal reasoning data set for insulator defect detection, and performs pretreatment, labeling and data enhancement on original inspection images; then adopts a LoRA parameter efficient fine-tuning method to adapt a visual language multi-modal large model to a field; embeds a visual tool interface in the fine-tuned model to construct a multi-modal reasoning framework with tool calling capability; further adopts a PPO reinforcement learning algorithm to jointly optimize the model by fusing a comprehensive reward function of format constraint, category correctness, boundary box matching and conditional tool calling; and finally inputs a to-be-detected image into the optimized model to output a structured detection result containing a defect type and a boundary box coordinate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of insulator defect detection technology, and in particular to an insulator defect detection method and system based on multimodal large model tool invocation. Background Technology

[0002] Insulators are critical components in substation equipment, responsible for electrical insulation, mechanical support, and conductor fixation. Their operational status directly affects the reliability of the power grid. Due to long-term exposure to complex environments such as high voltage, electric field stress, temperature and humidity variations, and contamination, insulators are prone to localized defects such as breakage, cracks, and gaps. These early defects are typically small in scale and have weak texture, and are easily affected by complex backgrounds, shooting angles, and changes in lighting. If not detected in time, they can lead to serious faults such as partial discharge and insulation breakdown. Therefore, research on automatic identification and localization of insulator defects based on substation inspection images is of great significance for improving inspection efficiency and ensuring power system safety.

[0003] Currently, the main technical approaches for insulator defect detection include traditional machine vision methods and deep learning-based target detection methods. Machine vision-based methods provide a technical foundation for early inspection image analysis, but feature extraction relies on manual design and has limited robustness. Infrared image-based methods utilize thermal features for fault identification, which has application value under specific imaging conditions, but is insensitive to early, small-scale defects. With the development of deep learning, researchers have applied two-stage detection networks (such as Faster R-CNN), single-stage detection networks (such as SSD and YOLO series), Visual Transformer (ViT), and multi-scale feature fusion and attention mechanisms to insulator defect identification, achieving certain results in complex environments and drone aerial photography scenarios. These studies indicate that feature enhancement targeting small targets and weak texture features on insulators is an important way to improve detection performance.

[0004] However, existing methods still have the following technical shortcomings: First, there is a lack of proactive and dynamic visual attention mechanisms. Traditional object detection models mostly rely on fixed network structures and one-time forward inference, making the detection process static. The model can only complete feature extraction and bounding box regression once on the entire image, unlike a human inspector who, after initial judgment, can "proactively" magnify, rotate, or adjust the perspective of suspected defect areas in the image for verification based on intermediate results. This "one-time" observation mode limits the model's ability to identify small and blurry defects.

[0005] Second, general-purpose multimodal large models are difficult to directly adapt to and stably execute in this task. In recent years, multimodal large models and intelligent agent tool invocation technology have provided new ideas for integrating visual perception and language reasoning. Some research has begun to apply them to power line inspection. However, when directly transferring general-purpose multimodal large models (such as the Qwen-VL series and GLM-4V series) to insulator defect detection tasks, there are obvious defects: on the one hand, the model output structure is unstable, often producing non-compliant natural language or incorrect JSON format, making automated parsing difficult; on the other hand, the model lacks the ability to finely locate small-scale defect regions, and when interfered with by complex backgrounds, the generated defect bounding boxes often show large offsets or omissions. The fundamental reason is that there are significant domain differences between the training data of general-purpose models and the equipment structure, defect morphology, and reasoning logic in power line inspection scenarios.

[0006] Third, existing data organization methods are insufficient to support the learning of the "observation-judgment-verification" reasoning process. General visual datasets or conventional "image-labeled bounding box" data formats can only train models for direct mapping, but cannot effectively express high-level reasoning strategies such as "when further observation is needed," "where to observe," and "how to use local observation information to correct the final judgment." This results in models lacking the data foundation to learn complex decision-making strategies.

[0007] In summary, how to endow insulator defect detection models with the ability to actively focus on local details, perform multi-round visual reasoning, and output stable and reliable structured results is a technical problem that urgently needs to be solved in this field. Summary of the Invention

[0008] The purpose of this invention is to provide an insulator defect detection method and system based on multimodal large model tools, aiming to solve the problems of low detection accuracy of small-scale and weak-texture defects in substation insulators, the inability of traditional detection models to actively focus on local details, and the unstable detection results of directly applying general multimodal large models in existing technologies. By endowing the model with active visual observation and dynamic reasoning capabilities, the accuracy and robustness of insulator defect detection in complex scenarios are significantly improved, facilitating integration with existing intelligent inspection systems.

[0009] To achieve the above objectives, this invention provides an insulator defect detection method based on multimodal large model tool invocation, the steps of which are as follows: S1. Construct a multimodal inference dataset for insulator defect detection, preprocess, annotate and augment the original inspection images to generate training samples containing task prompts, image inputs and structured labels; S2. The LoRA parameter efficient fine-tuning method is used to adapt the visual language multimodal large model to the domain, so that the model can master insulator defect discrimination, bounding box generation and standardized JSON output format. S3. Embed the visual tool interface into the fine-tuned model to build a multimodal inference framework with tool calling capabilities; S4. The PPO reinforcement learning algorithm is used to jointly optimize the large-scale visual language multimodal model. By integrating the comprehensive reward function of format constraints, category correctness, bounding box matching and conditional tool calls, the model is guided to learn a detection strategy that actively focuses on suspected defective regions. S5. Input the inspection image of the insulators of the substation to be inspected into the optimized visual language multimodal large model, and output the structured inspection results including the defect type and bounding box coordinates.

[0010] Preferably, step S1 specifically includes: S1.1 Image Acquisition and Preprocessing: Acquire multi-source insulator inspection images in real substation scenarios and divide them into training set, validation set and test set according to a preset ratio; S1.2 Boundary box coordinate correction: Let the image width and height be W and H respectively, and apply boundary constraints to the coordinates of all bounding boxes:

[0011] The corrected bounding box satisfies: Correct or remove annotation boxes that exceed the boundaries, have zero width and height, or have abnormal coordinate order; S1.3, Multimodal Inference Dataset Creation: The original images are uniformly scaled to a predetermined size; the LabelImg annotation tool is used to manually annotate the scaled images to generate an annotation file containing defect categories and bounding box coordinates; the images are data augmented and the bounding box coordinates are updated synchronously; the processed images, preset instructions, and JSON-formatted structured labels are combined into training samples; the JSON format includes at least the "defect_type" field and the "bboxes" field.

[0012] Preferably, step S2 specifically includes: The visual language multimodal large model consists of three parts: a visual encoder, a visual-language connector, and a large language model. The visual encoder is used to extract visual features from insulator inspection images, the visual-language connector is used to complete the dimensional mapping and semantic alignment of visual features and language features, and the large language model is used to generate text commands, tool call commands, and structured detection results. S2.1 LoRA Fine-tuning: Freeze the visual encoder and main parameters of the large visual-language multimodal model, and insert a low-rank adapter in the linear mapping layer of the model; let the original weight matrix be W, and the input feature be x, LoRA will transform the forward computation by... Become The weight update format is as follows: ; in, , , It is the rank of the low-rank decomposition, and , This is the scaling factor; during training, the original weights... Keep frozen, only update the low-rank matrix. and ; S2.2 Training Optimization: Suppose that the i-th training sample consists of an insulator image, text instructions and target output sequence. Supervised training is performed with the goal of minimizing the autoregressive cross-entropy loss. The loss function is summed by taking the negative logarithm of the sequence prediction probabilities of all samples.

[0013] Preferably, step S3 specifically includes: S3.1 Visual tool definition: Construct a set of visual tools, which includes at least an image local cropping and magnification tool; Tool invocation mechanism: The image local cropping and magnification tool receives the candidate region coordinates predicted by the visual language multimodal large model, crops the corresponding region from the original image and returns the local image. The local image returned by the tool is converted into a visual token and added back to the current context, so that the visual language multimodal large model can combine the original image and the local image to jointly determine the defect type and location in the subsequent generation process.

[0014] Preferably, step S4 specifically includes: S4.1 Markov Decision Process Modeling: The detection process is modeled as a Markov decision process. Let the state at step t consist of the original image, detection instructions, historical generated text, and historical tool observations. In the current state, the model selects an action according to the strategy. The action is a plain text token or a tool call instruction. S4.2 Comprehensive Reward Function Design: The output format, category judgment, bounding box localization, and tool invocation effectiveness are all included in the reward. The total reward consists of four parts: format reward, classification reward, bounding box localization reward, and tool invocation reward. The format reward is used to constrain the correctness of the JSON format of the model output. The classification reward is used to evaluate the consistency between the model's predicted insulator state category and the true category. The bounding box localization reward is used to evaluate the spatial matching degree between the predicted bounding box and the true labeled box. The tool invocation reward is used to evaluate the effectiveness of the model in calling visual tools.

[0015] Preferably, the bounding box localization reward in S4.2 Specifically, it includes: Bounding box localization reward in S4.2 Specifically, it includes: Intersection over Union (IoU) is used to measure the degree of overlap between predicted and ground truth bounding boxes. When the IoU is greater than a set threshold, the predicted bounding box is considered to be successfully matched with the ground truth bounding box. For cases with multiple defects in a single image, a greedy matching method is used to match the set of predicted bounding boxes with the set of ground truth bounding boxes. The number of successfully matched bounding boxes is counted, and the localization accuracy, recall rate, and average IoU are calculated. For normal samples, if the model outputs the category as "normal" and the bounding box set is empty, the localization result is considered correct. If the output is a non-empty bounding box, it is considered a false detection.

[0016] Preferably, the tool call reward in S4.2 adopts a conditional reward strategy, specifically including: The conditional tool reward is calculated by subtracting the invalid call penalty from the valid call reward. The valid call reward is triggered only when both the classification reward and the bounding box reward are positive. The invalid call penalty is proportional to the number of invalid tool calls. A tiered reward system is used during training: a negative reward is given for format errors, a lower positive reward is given for correct category but insufficient localization, and a higher reward is given when both the category and the bounding box are correct.

[0017] Preferably, S4 further includes: S4.3, PPO Policy Optimization: An actor-critic architecture is used for training. A LoRA-tuned visual-language multimodal large model is used as the initial actor model, and a visual-language multimodal large model of the same scale is used as the critic model for state value estimation. The ratio of action probabilities between the old policy and the current policy is set as the probability ratio. The PPO pruning policy objective is the expectation of the smaller of the product of the probability ratio and the advantage function and the product of the pruned probability ratio and the advantage function. The advantage function is jointly determined by the reward signal and the state value estimated by the critic. During training, the policy gradient is calculated only for the text tokens generated by the model and the tool call actions. The visual observations returned by the tool are processed by loss masking.

[0018] Preferably, step S5 specifically includes: S5.1 Preliminary judgment of the whole image: After receiving the inspection image of the insulator to be inspected, the visual language multimodal large model first extracts the features of the whole image through the visual encoder and makes a preliminary judgment. S5.2 Local Focus Verification: When the confidence level of the model in the preliminary judgment result is lower than the preset threshold, the visual language multimodal large model autonomous generation tool calls the instruction to output the coordinates of the candidate region, calls the image local cropping and magnification tool to locally magnify and observe the region, and the local image returned by the tool is added back to the context. The model combines the original image and the local image to perform verification judgment. S5.3 Structured Result Output: The visual language multimodal large model autonomously decides whether to make multiple rounds of tool calls based on the learned strategy until it obtains enough information to make a final judgment. The final output is kept in a unified JSON format. For defect samples, the output includes a list of defect types and bounding box coordinates. For normal samples, the output includes a list of normal types and empty bounding boxes.

[0019] An insulator defect detection system based on a multimodal large model tool is used to implement the above-described method, including: The dataset building module is used to collect and preprocess images of substation insulator inspections, and generate multimodal training samples containing task prompts, images, and structured labels. The model loading and fine-tuning module is used to load pre-trained visual language multimodal large models and use the LoRA parameter efficient fine-tuning method to perform domain adaptation training on them. The tool integration module is used to embed visual processing tools into the inference process of a large multimodal visual language model, enabling tool call command parsing and multimodal context management; The reinforcement learning optimization module is used to optimize the defect detection strategy and tool calling behavior of the visual language multimodal large model by combining the PPO algorithm with a multi-dimensional comprehensive reward function. The inference detection module receives the image to be detected, drives the visual language multimodal large model to perform multi-round inference with tool calls, and outputs the structured defect detection results.

[0020] According to specific embodiments provided by the present invention, the present invention discloses the following technical effects: 1. A high-quality dataset adapted for multi-turn inference was constructed and efficient domain adaptation was achieved. Breaking through the limitations of traditional image-labeled bounding box datasets, this approach supports the model in simultaneously learning defect recognition, bounding box localization, and multi-round inference logic through structured sample design. Combined with a unified coordinate system and bounding box correction mechanism, it enhances training stability. Employing the LoRA parameter efficient fine-tuning method, it achieves domain adaptation for general multimodal large models by training only a small number of low-rank adapter parameters. This reduces computational and memory overhead while resolving the issue of unstable structured output in general models.

[0021] 2. A detection strategy combining active visual observation and multi-dimensional collaborative optimization has been developed. A mechanism for calling image local cropping and magnification tools is introduced to simulate the overall observation-local focusing-detail confirmation reasoning process of manual inspection, giving the model the dynamic reasoning ability to actively focus on suspected defect areas. By integrating a comprehensive reward function that combines format constraints, classification correctness, bounding box matching and conditional tool calls, and combining PPO reinforcement learning to achieve joint optimization of detection strategies, the detection ability is significantly improved in small-scale, weak-texture and complex backgrounds while avoiding redundancy in blind tool calls.

[0022] 3. Possesses excellent engineering practicality and scenario scalability. This method achieves superior detection performance compared to general models with large parameters under conditions of small parameter count, reducing the hardware threshold for edge deployment; the standardized JSON output format can be seamlessly integrated with existing substation intelligent inspection systems, directly providing structured data support for inspection decisions; the technical framework has good versatility and can be quickly extended to other power equipment defect detection scenarios such as transformers and circuit breakers by replacing datasets and fine-tuning parameters.

[0023] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 This is a flowchart illustrating the creation of an insulator defect detection dataset in an embodiment of the present invention.

[0026] Figure 2 This is a training framework diagram for insulator defect detection in an embodiment of the present invention.

[0027] Figure 3 This is a schematic diagram of the structure of the visual language multimodal large model in an embodiment of the present invention.

[0028] Figure 4 This is a schematic diagram of an efficient fine-tuning scheme for multimodal large model parameters based on LoRA in an embodiment of the present invention.

[0029] Figure 5 This is a flowchart of the insulator defect detection multi-mode command fine-tuning data construction process in an embodiment of the present invention.

[0030] Figure 6 This is a flowchart of the reinforcement learning tool invocation strategy optimization based on PPO in an embodiment of the present invention.

[0031] Figure 7 This is a comparison chart of the average number of tool calls under different tool reward settings in embodiments of the present invention.

[0032] Figure 8 This is a comparison chart of the detection results of different multimodal large models on typical insulator defect samples in the embodiments of the present invention.

[0033] Figure 9 This is an example diagram of the intermediate inference process of the model in an embodiment of the present invention. Detailed Implementation

[0034] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0035] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0036] Example This embodiment provides a substation insulator defect detection method based on multimodal large model tool invocation, the overall framework of which is as follows: Figure 2 As shown, it mainly includes three parts: the construction of the insulator multimodal inference dataset, LoRA-based supervised fine-tuning, and PPO-based reinforcement learning tool invocation strategy optimization.

[0037] S1. Construct a multimodal inference dataset for insulator defect detection, preprocess, annotate, and augment the original inspection images to generate training samples containing task prompts, image inputs, and structured labels.

[0038] S1.1 Image Acquisition and Preprocessing The insulator images used in this embodiment are from real substation inspection scenarios. The acquisition methods include drone and manual inspection, covering different distances, angles, lighting conditions, and background conditions. A total of 1694 images were acquired, including 1212 defect images and 482 normal images, which were divided into training, validation, and test sets in a 7:2:1 ratio.

[0039] S1.2 Boundary box coordinate correction To address the training instability caused by differences in pixel coordinate ranges, a bounding box coordinate correction mechanism is introduced during the data construction phase. Let the image width and height be... and Apply boundary constraints to the coordinates of all annotation boxes:

[0040] And ensure that the corrected bounding box satisfies: .

[0041] For bounding boxes that exceed the image boundaries, have zero width and height, have abnormal coordinate order, or are invalid after correction, they should be corrected or removed; a unified coordinate system should be used in image scaling, cropping, JSON label generation, and IoU calculation to reduce label errors and misjudgments of rewards caused by pixel-level offsets.

[0042] For tool usage training in the PPO phase, samples with minor defects and where the entire image is difficult to judge directly but local features are obvious are prioritized, enabling the model to learn a "holistic observation - local verification" reasoning strategy. Samples with excessively large defects, severely blurred images, or ambiguous annotations are given lower priority in the reinforcement learning phase. Simultaneously, all samples are output in a JSON-only format, retaining only the `defect_type` and `bboxes` fields to facilitate automatic parsing and calculation of classification, localization, and structured output metrics. For multi-defect images, multiple bounding boxes are organized into a unified list format to support multi-defect detection in a single image.

[0043] S1.3, Multimodal Inference Dataset Creation To meet the training requirements for image understanding, defect localization, and structured representation, the original inspection images, manually labeled images, task prompts, and standard answers are uniformly organized into multimodal reasoning samples. The overall production process is as follows: Figure 1 As shown.

[0044] (1) Image scale unification: The original image is preprocessed uniformly, and the image size is adjusted to 1960×1960 pixels to ensure consistent input scale, reduce training complexity, and improve the model's learning efficiency for local defect regions of insulators. Let the original image size be... The scaled image size is Then any point in the original graph ( The coordinates after scaling can be expressed as:

[0045] In the formula, ( () represents the corresponding coordinates in the scaled image; in this embodiment, it is taken as... By using the above-mentioned scale unification process, the consistency of input formats among different samples can be enhanced while preserving the main structural information and defect texture features, thereby providing a unified data foundation for subsequent model training.

[0046] (2) Manual annotation: LabelImg was used to manually annotate the ceramic and composite insulator samples. Since the annotation was done directly on the scaled image, the bounding box coordinates were consistent with the training input, reducing the error caused by multiple coordinate mappings.

[0047] (3) Data Augmentation: To improve generalization ability, the sample was expanded using augmentation methods such as rotation, mirroring, noise injection, and brightness adjustment to simulate common viewpoint changes, imaging noise, and illumination fluctuations in inspections, thereby improving the model's robustness to insulator defects in complex scenarios. Let the original image be... Data augmentation operations are denoted as ,in Let the enhancement parameters be represented as follows:

[0048] Accordingly, the augmented sample can be represented as:

[0049] in, This represents the label information corresponding to the augmentation result. For data augmentation operations involving geometric transformations, the corresponding bounding box positions need to be updated synchronously; for non-geometric augmentation operations such as blurring, noise reduction, and brightness adjustment, the label information remains unchanged. By using a combination of augmentation strategies, the coverage and distribution richness of training samples can be effectively improved.

[0050] After image preprocessing, manual annotation, and data augmentation, the samples are further organized into a structured format suitable for multimodal inference training. Let the... Each sample is represented as:

[0051] in, This indicates a task prompt message. Indicates the input image. Represents structured tags, This indicates additional information such as image name, sample number, and data partitioning. Structured tags. This can be further expressed as:

[0052] in, Indicates the sample category, This represents the set of defect bounding boxes. When no defects exist in the sample... When there is one or more defects in the sample, This design can support both normal sample discrimination and single-defect and multi-defect localization, thus better adapting to the complex and variable defect distribution characteristics of real-world inspection scenarios.

[0053] S2. The LoRA parameter efficient fine-tuning method is used to adapt the visual language multimodal large model to the domain, enabling the model to master insulator defect discrimination, bounding box generation and standardized JSON output format.

[0054] The Vision-Language Multimodal Large Model (VLM) consists of three core modules: a Vision Encoder, a Vision-Language Connector, and a Large Language Model (LLM). Its overall architecture is as follows: Figure 3 As shown, the model first encodes the visual features of the input inspection image, then completes the semantic alignment between visual and linguistic features through a vision-language connector, and finally inputs it into a large language model to complete cross-modal inference and generate structured detection results. Let the input inspection image be... The text detection command is The model ultimately outputs the structured detection results as follows: The entire visual language multimodal big model can then be represented as:

[0055] in, Indicates a visual encoder; This indicates a visual-language connector; Indicates text embedding; This represents a large language model. The entire model completes cross-modal mapping from image and text input to structured detection results output.

[0056] (1) Visual encoder The visual encoder employs a Vision Transformer (ViT) structure, its function being to encode the input inspection image into a sequence of visual features. Let the input image be... First, the image is divided into The size is Image patches.

[0057] No. The initial embedding is obtained by linear mapping of the image patches:

[0058] in The Patch projection matrix; For position encoding.

[0059] Therefore, the initial visual token sequence is represented as: .

[0060] Then enter by A visual encoder composed of layers of Transformers.

[0061] No. The layer Transformer calculation process is represented as follows:

[0062]

[0063] in, This indicates the focus of the bulls; Indicates a feedforward network; This indicates root mean square normalization.

[0064] To reduce computational complexity, the Qwen2.5-VL visual encoder employs an alternating stacking structure of window attention and global attention, enabling the model to maintain its ability to model local details while establishing global connections across regions.

[0065] The feedforward network uses the SwiGLU activation function:

[0066]

[0067] The final visual feature sequence is obtained: ; Each visual token contains local texture and global semantic information for the corresponding image region.

[0068] (2) Visual-Language Connector Since the output dimension of the visual encoder differs from the hidden space of the language model, a vision-language connector is needed to complete cross-modal feature mapping.

[0069] Qwen2.5-VL employs a Patch Merge strategy to spatially aggregate adjacent 2×2 visual tokens, and then projects them to the language model's hidden space through two layers of MLP.

[0070] Let the four adjacent visual tokens be:

[0071] The corresponding new visual token is represented as follows:

[0072] After Patch Merge, the number of visual tokens is reduced to approximately one-quarter of its original value; therefore, the connector output is represented as: ,in ; The connector not only completes the mapping of visual feature dimensions, but also realizes the unified representation of visual spatial information to the linguistic semantic space, providing a unified input for subsequent cross-modal reasoning.

[0073] (3) Large language model Text detection instructions are first mapped to a text token sequence. Then the visual token and the text token are concatenated to obtain... , and input Transformer Decoder.

[0074] No. Layer attention calculation is represented as:

[0075] Corresponding self-attention calculation:

[0076] in As a causal mask, it ensures that only historical tokens are used during the generation process.

[0077] The Transformer layer is updated to:

[0078] After all Transformer layers, the first... The predicted probability for each token is:

[0079] Therefore, the entire output sequence satisfies:

[0080] The final result is a JSON-formatted detection result, including the defect category and bounding box coordinates.

[0081] (4) Monitor and fine-tune the objectives To adapt the basic visual language model to the insulator defect detection task, this invention employs the LoRA parameter efficient fine-tuning method, which trains only the low-rank adapter parameters while freezing the main parameters of the pre-trained model.

[0082] Supervised fine-tuning employs autoregressive cross-entropy loss:

[0083] LoRA does not directly update the original weight matrix. And learning low-rank increments .in, .

[0084] Therefore, the linear layer is composed of Updated to ,in, The scaling factor is used; the original parameters are kept frozen, and only the low-rank matrix is ​​optimized. and .

[0085] After supervised fine-tuning, the model can establish a mapping relationship between insulator images, defect categories, and bounding box coordinates, providing an initial model with domain knowledge for subsequent optimization of tool invocation strategies based on reinforcement learning.

[0086] (5) PPO-based reinforcement learning optimization To further enhance the model's tool-calling capability in small-scale insulator defect detection tasks, this invention introduces the Proximal Policy Optimization (PPO) algorithm for reinforcement learning training of the model, based on supervised fine-tuning. The model models the visual tool-calling process as a sequential decision problem, guiding the model to learn when to call visual tools at appropriate times through a reward function, and optimizing the final detection results.

[0087] Let the reinforcement learning process be represented as a Markov Decision Process (MDP):

[0088] in, Representing the state space, Represents the action space. Represents the state transition probability. Represents the reward function, This is the discount factor.

[0089] In the Step, the model state is defined as:

[0090] in, This indicates the input inspection image. Indicates a detection command. This represents the currently generated text sequence. This represents the local observation information returned by the vision tool. Based on the current state, the model outputs an action through the policy network: ;in, The parameter is The strategy model includes actions that encompass both text token generation and visual tool invocation decisions, such as whether to invoke image cropping tools, rotation tools, and corresponding region selections.

[0091] To simultaneously optimize detection results and tool invocation behavior, this invention constructs a joint reward function:

[0092] in, Used to constrain the model output to conform to a predefined JSON format; Used to evaluate the accuracy of defect category prediction; Rewards are given based on the degree of matching between the predicted bounding box and the true bounding box; This is used to evaluate the visual tool invocation strategy. When the tool invocation can effectively improve the detection quality, a positive reward is given, and otherwise a penalty is given, thereby guiding the model to learn a reasonable tool use strategy.

[0093] The PPO algorithm improves the stability of reinforcement learning training by limiting the policy update magnitude. Its optimization objective is defined as: ;in,

[0094] This represents the probability ratio between the old and new strategies. Represents the dominance function. This is the policy truncation coefficient, used to limit the magnitude of a single policy update and ensure stable convergence during training.

[0095] After PPO reinforcement learning optimization, the model can autonomously decide whether to invoke visual tools based on the current detection results, and continuously correct defect category judgment and bounding box localization by combining local observations, achieving joint optimization of detection results and tool invocation strategy. Ultimately, the optimization objective of the overall model of this invention can be expressed as:

[0096] in, This represents the model parameters optimized by reinforcement learning. By maximizing the cumulative reward, the model not only improves the accuracy of insulator defect detection but also learns more efficient and stable vision tool invocation strategies, further enhancing its ability to detect small-scale defects in complex inspection scenarios.

[0097] S2.1 Efficient Parameter Fine-Tuning Based on LoRA This embodiment employs the LoRA parameter efficient fine-tuning method to perform domain adaptation on the basic multimodal large model. LoRA fine-tuning is used to complete domain alignment, enabling the general multimodal large model to master the insulator defect image features and a unified structured output format, and providing a stable initial strategy for subsequent PPO training. The LoRA fine-tuning structure is as follows: Figure 4 As shown.

[0098] (1) Fine-tuning the data format The fine-tuning data consists of insulator images, detection commands, and JSON annotation results, such as... Figure 5 As shown. Defect samples output "damage" and bounding box coordinates, while normal samples output "normal" and an empty bounding box. The target output format is defined as follows:

[0099] Here, `defect_type` represents the insulator state category, and `bboxes` represents the set of bounding boxes for the defect region. Each bounding box is defined by the coordinates of its top-left and bottom-right corners. This means that, through this data organization method, the original object detection labels are transformed into a structured text sequence that can be learned by a multimodal large model.

[0100] (2) LoRA fine-tuning principle The core idea of ​​LoRA fine-tuning is to introduce low-rank trainable matrices into the linear mapping layers of the model while freezing the main parameters of the pre-trained model. Let the original weight matrix be... The input features are The original linear transformation can be expressed as: LoRA does not update directly. Instead, it learns a low-rank increment. This makes the forward computation become: ; in, It is decomposed into the product of two low-rank matrices: ; Therefore, the weight update form after adding LoRA is: ; In the formula, , , It is the rank of the low-rank decomposition, and , This is the scaling factor. During training, the original weights... Keep frozen, only update the low-rank matrix. and This approach enables model adaptation with fewer trainable parameters, reducing the memory and computational overhead of fine-tuning large models.

[0101] The main parameters of the visual encoder and the base large language model are kept frozen, with only a low-rank adapter added to the linear layers and a small number of task-related parameters trained. The aforementioned low-rank matrix... and We learn the incremental parameters related to the task together, and then train the model using supervised loss.

[0102] S2.2 Training Configuration and Optimization Objectives Let the first The training samples consist of insulator images. Text instructions and target output sequence Composition. The LoRA fine-tuning stage uses autoregressive cross-entropy loss for supervised training, with the optimization objective being:

[0103] in, This represents the frozen basic multimodal large model parameters. This represents the trainable parameters introduced by LoRA. Indicates the first in the target sequence Each token. This loss function enables the model to progressively learn manually annotated JSON outputs given an image and detection instructions.

[0104] In the specific implementation, the Qwen2.5-VL-7B-Instruct is used as the base model, with images and detection instructions as inputs and JSON annotations as supervision signals. Through this training stage, the model establishes a mapping relationship between insulator image features, state categories, and bounding box coordinates.

[0105] LoRA fine-tuning achieves basic detection capabilities with low memory overhead and provides a good initial model for the PPO stage, allowing subsequent optimizations to focus on output quality, localization accuracy, and tool calling strategies.

[0106] S3, Building Tools to Enhance Reasoning Frameworks By embedding a visual tool interface into the fine-tuned model, a multimodal inference framework with tool call capabilities is constructed.

[0107] During the tool-enhanced inspection process, the model receives the original insulator inspection images. and detection instructions Then, based on the current visual information and the generated content, two types of actions can be selected: one is to continue generating text content, and the other is to invoke a visual tool. In this embodiment, the visual tool used is a local image magnification tool, whose input is the coordinates of the candidate region predicted by the model. The output is a cropped image of the corresponding region. The local image returned by the tool is reintegrated into the current context, enabling the model to combine the original image and the local image in subsequent generation processes to determine the type and location of defects. This process can be represented as:

[0108]

[0109] in, This tool represents a zoom-in area of ​​an image. This represents the local visual observation obtained after the t-th tool call. This is different from directly based on the whole... Figure 1 Compared to the one-time output, this interactive process can provide the model with higher resolution local details, and is especially suitable for confirming small-scale defects such as insulator self-explosion, breakage, and missing parts.

[0110] S4: Joint optimization using PPO reinforcement learning The PPO reinforcement learning algorithm is used for joint optimization of the model. By fusing a comprehensive reward function that integrates format constraints, class correctness, bounding box matching, and conditional tool invocation, the model is guided to learn a detection strategy that actively focuses on suspected defect regions. The optimization process of the PPO-based reinforcement learning tool invocation strategy is as follows: Figure 6 As shown.

[0111] After fine-tuning with LoRA instructions, the LoRA-SFT model already possesses basic detection capabilities, but it may still miss small defects, cause bounding box offsets, or have unstable formats in complex backgrounds. To address this, PPO reinforcement learning is further introduced, modeling the detection process as a multi-round decision-making process with visual tool calls, enabling the model to perform local magnification and verification of suspected areas when needed.

[0112] S4.1 Markov Decision Process Modeling: The above process is formalized as a Markov decision process. Let the first... The state of the step is It consists of the original image, detection instructions, historical generated text, and historical tool observations: ;in, This represents the text sequence generated by the model up to the current time step. This represents the set of local image observations obtained by calling the tool.

[0113] In state The model follows the strategy. Select Action : ;action It can be a plain text token or a tool invocation command. When the model selects the tool invocation action, it generates candidate bounding box coordinates and obtains the corresponding local image; when the model believes that the existing information is sufficient to make a judgment, it generates the final detection result.

[0114] A complete detection trajectory can be represented as: ;in, Y represents the visual observation or environmental feedback returned by the tool. The final output is in JSON format consistent with the LoRA-SFT stage, containing two fields: defect_type and bboxes. The defect_type can be either damage or normal, and bboxes is an array of bounding box coordinates.

[0115] This design ensures that the goals of the supervised fine-tuning and reinforcement learning phases are consistent, namely, always optimizing the model around "state judgment + defect localization".

[0116] S4.2 Comprehensive Reward Function Design: The key to the PPO stage lies in the reward function, which incorporates output format, class determination, bounding box localization, and tool call effectiveness into the reward, making the model optimization objective closer to the actual detection quality.

[0117] in, Used to constrain whether the model output meets the JSON format requirements; Used to determine whether the output category matches the actual label; Used to evaluate the spatial matching degree between predicted bounding boxes and ground truth bounding boxes; This is used to evaluate the effectiveness of tool calls. Through this reward design, the model's optimization objective shifts from simply generating standard answers to directly optimizing the final quality of the detection task.

[0118] The format reward requires the output to be parsable JSON and to contain only the defect_type and bboxes fields to reduce irrelevant natural language and structural errors.

[0119] For categorical rewards, the judgment is based on the consistency between the predicted category and the actual category. Let the predicted category be... The real category is The correctness of the classification can then be expressed as:

[0120] Among them, categories This includes classifying samples as "normal" and "damage." If the model misclassifies a defective sample as "normal" or a normal sample as "damage," the classification reward is reduced. This item primarily enhances the model's ability to make global judgments about the insulator's state.

[0121] For bounding box localization rewards, the Intersection over Union (IoU) ratio is used to measure the overlap between the predicted and ground truth bounding boxes. Let the predicted bounding box be... The real frame is Then the IoU between the two is defined as:

[0122] when Greater than the set threshold When the predicted bounding box matches the ground truth bounding box, a greedy matching method is used for the set of predicted bounding boxes in a single image with multiple defects. With the set of real boxes Perform matching, count the number M of successfully matched bounding boxes, and calculate the localization precision, recall rate, and average IoU:

[0123]

[0124]

[0125] in, This indicates the proportion of correctly matched items in the prediction box. This indicates the percentage of actual defects that were detected. This indicates the average overlap of the matching boxes. For normal samples, if the model outputs a "normal" category and the bounding box set is empty, the localization result is considered correct; if the output is a non-empty bounding box, it is considered a false detection.

[0126] Tool call rewards are used to constrain the model's reasonable use of local observation tools. Since defect regions are usually small, local magnification helps confirm suspected defects; however, without conditional encouragement for tool calls, it may lead to redundant pruning and inference redundancy. Therefore, a conditional tool call reward is adopted: a reward is given only when the final detection is correct and the called region is valid, and out-of-bounds, excessively small, repeated, or irrelevant calls are penalized.

[0127]

[0128] in, For indicator functions, Indicates a reward for effective tool usage. Indicates the number of invalid tool calls. This is the penalty coefficient for invalid calls. This design prevents the model from blindly calling tools in order to obtain rewards, ensuring that tool calls are consistent with the final detection quality.

[0129] A tiered reward system is used during training: negative rewards are given for format errors, lower positive rewards are given for correct categories but insufficient localization, and higher rewards are given for both correct categories and bounding boxes, guiding the model to gradually improve from "valid output" to "accurate localization".

[0130] S4.3, PPO Strategy Optimization PPO ensures training stability by limiting the range of variation between the old and new strategies. Let the old strategy be... The current strategy is Then the first The probability ratio of each step action is:

[0131] The goal of PPO's trimming strategy is:

[0132] in, For the dominant function, Here, represents the pruning factor. The advantage function is determined by both the reward signal and the state value estimated by the critic:

[0133] in, Indicates the critic's state The value estimation is performed by the actor, which updates the generation policy based on the advantage function, while the critic provides a more stable value estimate by fitting the reward. Since the image observations returned by the tool are not tokens actively generated by the model, a loss mask is applied to the tool observation part during training, and the policy gradient is only calculated for the model-generated text tokens and tool invocation actions. This ensures that PPO optimization focuses on the model's own decision-making behavior, rather than meaninglessly fitting external observations.

[0134] In the PPO stage, the LoRA-SFT model is used as the initial actor. Through multiple rounds of rollout and policy updates, the model gradually learns to select more reasonable observation regions in complex inspection images and outputs more stable structured detection results.

[0135] Overall, LoRA-SFT is responsible for establishing basic detection capabilities, while PPO further incorporates format, category, location, and tool calls into a unified optimization, thereby improving the detection stability of small-scale defects and complex backgrounds.

[0136] S5. Input the inspection image of the insulators of the substation to be inspected into the optimized visual language multimodal large model, and output the structured inspection results including the defect type and bounding box coordinates.

[0137] S5.1 Preliminary judgment of the whole image: After receiving the inspection image of the insulator to be inspected, the visual language multimodal large model first extracts the features of the whole image through the visual encoder and makes a preliminary judgment. S5.2 Local Focus Verification: When the model's confidence level in the preliminary judgment result is lower than a preset threshold, the visual language multimodal large model autonomous generation tool invokes instructions to output the coordinates of the candidate region. The image local cropping and magnification tool is invoked to magnify and observe the area. The local image returned by the tool is added back into the context, and the model combines the original image and the local image for verification and judgment. S5.3 Structured Output: The visual language multimodal large model autonomously decides whether to perform multiple rounds of tool calls based on the learned strategy until sufficient information is obtained to make a final judgment. The final output is kept in a unified JSON format: Defect Sample Output:

[0138] Normal sample output: .

[0139] II. Experimental Results and Analysis 1. Experimental Environment and Parameter Configuration The experiments in this embodiment were conducted in a Linux environment, using four NVIDIA H100 PCIe GPUs as the computing platform. The supervised fine-tuning phase employed LLaMA-Factory for LoRA training, the reinforcement learning phase utilized the PPO training framework, and the inference phase combined Transformers and vLLM to accelerate model loading and generation. All comparative experiments used the same dataset, output parsing method, and evaluation metrics.

[0140] This embodiment uses Qwen2.5-VL-7B-Instruct as the base model, and the training process includes two stages: LoRA supervised fine-tuning and PPO reinforcement learning. The training samples consist of insulator images, detection instructions, and JSON annotations. Defect images output "damage" and bounding boxes, while normal images output "normal" and empty bounding boxes.

[0141] During the LoRA fine-tuning phase, the LoRA rank r is set to 8, and the scaling factor is... The settings are set to 16, Dropout is set to 0.05, the optimizer is AdamW, and the learning rate is set to... The learning rate scheduling method uses cosine annealing. Considering the high memory consumption of multimodal image input, the batch size per GPU is set to 4, and the gradient accumulation steps are set to 4, resulting in an equivalent global batch size of 64 on four H100PCIe GPUs. The training epochs are set to 100, with evaluation on the validation set every fixed number of steps. The optimal LoRA-SFT model is saved based on the validation set loss and the stability of the structured output. In this stage, by freezing the main parameters of the base model and training only the LoRA low-rank adapter, the model achieves insulator defect category judgment and bounding box generation capabilities with lower training overhead.

[0142] In the PPO reinforcement learning phase, the LoRA-SFT-adapted Qwen2.5-VL-7B model is used as the initial actor model, and a 7B-scale model is used as the critic for state value estimation. The actor is responsible for generating defect detection results or tool invocation actions based on the insulator image and detection instructions, while the critic is responsible for estimating the current state value and assisting in the calculation of the advantage function. The learning rates for both the actor and critic in the PPO phase are set to [value missing]. During training, BF16 mixed precision and a multi-GPU parallel strategy were employed. To adapt to high-resolution insulator inspection images and multi-round tool calls, the maximum sequence length was set to 14336, with the maximum prompt length set to 12288 and the maximum response length set to 2048. Output format constraints were enabled during training to restrict the generation of irrelevant natural language, inference labels, and non-target fields, ensuring that the model's final output remained a parsable JSON structure.

[0143] 2. Analysis of the impact of tool call rewards on model detection performance To verify the impact of tool call rewards during the Proof-of-Detection (PPO) phase on model detection behavior, three ablation experiments were conducted while maintaining consistency in the visual tool interface, training data, and PPO parameters: no tool reward, unconditional tool reward, and conditional tool reward. In the no-tool reward experiment, optimization was based solely on output format, class determination, and bounding box localization quality. The unconditional tool reward awarded a reward whenever the model called a visual tool. The conditional tool reward experiment required that tool calls be associated with the final correct detection result and imposed constraints on out-of-bounds, duplicate, or irrelevant region calls. This experiment aimed to analyze whether tool rewards could guide the model to form an effective local observation strategy, rather than simply increasing the number of tool calls.

[0144] Average number of tool calls under different reward settings, such as Figure 7 As shown, the unconditional tool reward curve drops significantly in the later stages of training, indicating that in the absence of explicit incentives, the model tends to reduce local observations and directly output detection results based on the entire image. The unconditional tool reward maintains a certain calling frequency, but the overall change is relatively gradual, indicating that the model forms more of a fixed calling habit, making it difficult to determine whether the tool calls truly help in defect identification. In contrast, the conditional tool reward gradually increases and tends to stabilize in the later stages of training, indicating that the model gradually learns to call tools such as local cropping and magnification when needed to verify suspected defect areas.

[0145] This change aligns with the characteristics of insulator defect detection tasks. Defects such as breaks and gaps in inspection images are typically small in scale and easily affected by the background, adjacent insulators, and the shooting angle. Appropriate tool calls can provide the model with local detail information, helping to improve category judgment and bounding box localization; however, if the reward only encourages the calling behavior itself, it may introduce duplicate observations and invalid clipping. Therefore, the key to tool rewards is to constrain the model to "use the tool effectively," rather than simply pursuing a higher number of calls.

[0146] Table 1 presents the detection results under different tool reward settings. It can be seen that conditional tool rewards achieve the best performance in Acc, Precision, and Recall, indicating that binding tool calls to the final detection quality can more effectively improve model performance. Without tool rewards, the model makes fewer tool calls in the later stages, resulting in insufficient local observation of small-scale defects and a relatively low recall rate. While unconditional tool rewards show some improvement compared to no tool rewards, the performance gain is limited because it does not constrain whether tool calls improve classification and localization results.

[0147] Table 1. Results of ablation experiments using the tool call reward mechanism.

[0148] comprehensive Figure 7 As shown in Table 1, tool usage is practically significant for insulator defect detection, but its effectiveness depends on the reward design. Conditional tool rewards can encourage the model to actively focus on suspected defect areas in complex samples and transform local observations into more accurate category judgments and bounding box localizations; at the same time, they prevent the model from blindly using tools to obtain rewards. This indicates that jointly constraining tool usage behavior with the quality of detection results helps improve the detection stability in small-scale, weakly textured defect scenarios.

[0149] 3. Defect detection and judgment criteria To evaluate the performance of the visual language multimodal large model in insulator defect detection, this embodiment uses defect detection accuracy as the main indicator. Since the model output is structured text, the JSON format must be parsed before calculating the accuracy, and the accuracy of the sample detection must be determined by combining the category judgment and bounding box localization results.

[0150] The model output should uniformly include two fields: `defect_type` and `bboxes`. Normal samples should output "normal" with empty bounding boxes; defective samples should output "damage" along with the coordinates of the defect area. If the output cannot be parsed, has missing fields, or does not meet the format requirements, the sample is considered incorrect.

[0151] For the Zhang test images, assuming the true category is... The model predicts the category as .like =normal, then only if the model prediction result satisfies: At that time, the sample was judged as correct. Among them, This represents the set of bounding boxes predicted by the model. If the model misclassifies a normal sample as a defect, or outputs a non-empty defect box, it is considered an error.

[0152] like If the model outputs "damage", then it not only needs to correctly output the defect category, but also needs to provide bounding boxes that match the real defect regions. Let the set of real defect boxes be... The predicted defect box set is For the predicted bounding box and real frame Its intersection-union ratio is defined as:

[0153] When the predicted category satisfies =damage, and the maximum IoU between the predicted bounding box and the ground truth bounding box is not lower than the set threshold. At that time, the defect is considered to have been correctly detected.

[0154] Threshold in this embodiment Set to 0.5. For images with multiple defect boxes, the matching result between the predicted box and the ground truth box is used for judgment. When the model can correctly identify the defect category and at least complete the effective localization of the main defect region, the sample is judged as a correct detection; if the model classifies incorrectly, does not output a valid bounding box, or the overlap between the predicted box and the ground truth box is less than the threshold, it is judged as an error.

[0155] Based on the above discrimination rules, this embodiment will... The test results for each test sample are recorded as follows: :

[0156] The defect detection accuracy on the test set is defined as:

[0157] in, This represents the total number of test samples. This metric considers whether the model output format, normal / defect category judgment, and defect area location meet the requirements, and can reflect the comprehensive detection capability of the multimodal large model in insulator defect detection tasks.

[0158] To ensure fair comparison, all models use the same input image, detection instruction template, and discrimination rules. For general multimodal large models that cannot directly output bounding boxes, they are still parsed using a unified JSON format; if the format is incorrect or localization cannot be completed, they are counted as erroneous samples.

[0159] 4. Insulator Defect Detection Results and Analysis To verify the detection performance of the method in a real substation inspection scenario, this embodiment compares the proposed model with several general multimodal large models, including Qwen3-VL-Plus, Qwen-VL-Max, Qwen2.5-VL-7B-Instruct, Qwen2.5-VL-32B-Instruct, Qwen3-VL-235B-A22B-Instruct, and GLM-4.6V-Flash. The results are shown in Table 2. The model in this embodiment achieves an accuracy of 82.94% at a 7B parameter scale, higher than all the compared models. The best-performing competing model is Qwen3-VL-Plus, with an accuracy of 82.49%, slightly lower than the model in this embodiment. The accuracy of the other general models ranges from 62.21% to 71.43%. The model in this embodiment uses only a 7B parameter scale, which is significantly smaller than some of the general multimodal large models compared. However, after two-stage optimization training tailored to the specific task, its detection accuracy is still higher than that of larger-scale models such as 32B and 235B. This result indicates that in domain tasks such as insulator defect detection, parameter scale is not the only factor affecting model performance, and task-specific adaptation training is more helpful in improving detection results.

[0160] Compared to the base model Qwen2.5-VL-7B-Instruct, the accuracy of this embodiment model improved from 65.90% to 82.94%, damage accuracy from 55.63% to 77.46%, and normal accuracy from 85.33% to 93.33%. Compared to Qwen3-VL-Plus, the damage accuracy of this embodiment model is slightly lower than its 80.28%, but the normal accuracy is higher than its 86.67%, achieving better overall accuracy results. This indicates that the model of this embodiment has better comprehensive performance in both defect sample identification and normal sample judgment.

[0161] Table 2 Comparison of detection results for different multimodal large models

[0162] To further compare the defect localization capabilities of different models, typical insulator defect samples were selected for visual comparison of the detection results. The results are as follows: Figure 8 As shown in the figure, although Qwen-VL-Max identifies a defect in the image, its predicted bounding box is offset to the right side of the non-main insulator area, failing to accurately locate the actual defect location. The 7B base model also identifies a defect, but the predicted bounding box deviates significantly from the actual damaged area. This indicates that general multimodal large models without domain training are easily affected by background insulators and adjacent equipment structures in fine-grained defect localization. In contrast, the model in this embodiment can correctly identify the presence of insulator defects in the image and locate the predicted bounding box to the actual damaged area in the lower part of the main insulator, demonstrating a more stable defect localization capability.

[0163] In addition to the final detection results, this embodiment further demonstrates the model's intermediate inference process and tool invocation behavior, such as... Figure 9 As shown, the model first determines whether defects exist based on the overall image information, then uses a local zoom-in tool to inspect the edges of the main insulator and the right-side glass insulator area. After eliminating background structural interference, it focuses on the damaged location in the lower middle part of the main insulator and outputs the bounding box. This process demonstrates that the tool call enables the model to perform multiple rounds of reasoning and judgment, making fuller use of the original image and local image information, which helps improve the accuracy of defect judgment and location.

[0164] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. An insulator defect detection method based on multimodal large model tool invocation, characterized in that, The steps are as follows: S1. Construct a multimodal inference dataset for insulator defect detection, preprocess, annotate and augment the original inspection images to generate training samples containing task prompts, image inputs and structured labels; S2. The LoRA parameter efficient fine-tuning method is used to adapt the visual language multimodal large model to the domain, so that the model can master insulator defect discrimination, bounding box generation and standardized JSON output format. S3. Embed a visual tool interface into the fine-tuned model to build a multimodal inference framework with tool call capabilities; S4. The PPO reinforcement learning algorithm is used to jointly optimize the large-scale visual language multimodal model. By integrating the comprehensive reward function of format constraints, category correctness, bounding box matching and conditional tool calls, the model is guided to learn a detection strategy that actively focuses on suspected defective regions. S5. Input the inspection image of the insulators of the substation to be inspected into the optimized visual language multimodal large model, and output the structured inspection results including the defect type and bounding box coordinates.

2. The insulator defect detection method based on multimodal large model tool invocation according to claim 1, characterized in that, Step S1 specifically includes: S1.1 Image Acquisition and Preprocessing: Acquire multi-source insulator inspection images in real substation scenarios and divide them into training set, validation set and test set according to a preset ratio; S1.2 Boundary box coordinate correction: Let the image width and height be W and H respectively, and apply boundary constraints to the coordinates of all bounding boxes: ; The corrected bounding box satisfies: Correct or remove annotation boxes that exceed the boundaries, have zero width and height, or have abnormal coordinate order; S1.3, Multimodal Inference Dataset Creation: The original images are uniformly scaled to a predetermined size; the LabelImg annotation tool is used to manually annotate the scaled images to generate an annotation file containing defect categories and bounding box coordinates; the images are data augmented and the bounding box coordinates are updated synchronously; the processed images, preset instructions, and JSON-formatted structured labels are combined into training samples; the JSON format includes at least the "defect_type" field and the "bboxes" field.

3. The insulator defect detection method based on multimodal large model tool invocation according to claim 1, characterized in that, Step S2 specifically includes: The visual language multimodal large model consists of three parts: a visual encoder, a visual-language connector, and a large language model. The visual encoder is used to extract visual features from insulator inspection images, the visual-language connector is used to complete the dimensional mapping and semantic alignment of visual features and language features, and the large language model is used to generate text commands, tool call commands, and structured detection results. S2.1 LoRA Fine-tuning: Freeze the main parameters of the visual encoder and the large language model, and insert a low-rank adapter in the linear mapping layer of the model; let the original weight matrix be W, and the input feature be x, LoRA will transform the forward computation by... Become The weight update format is as follows: ; in, , , It is the rank of the low-rank decomposition, and , This is the scaling factor; during training, the original weights... Keep frozen, only update the low-rank matrix. and ; S2.2 Training Optimization: Suppose that the i-th training sample consists of an insulator image, text instructions and target output sequence. Supervised training is performed with the goal of minimizing the autoregressive cross-entropy loss. The loss function is summed by taking the negative logarithm of the sequence prediction probabilities of all samples.

4. The insulator defect detection method based on multimodal large model tool invocation according to claim 1, characterized in that, Step S3 specifically includes: S3.1 Visual tool definition: Construct a set of visual tools, which includes at least an image local cropping and magnification tool; S3.2 Tool Invocation Mechanism: The image local cropping and magnification tool receives the candidate region coordinates predicted by the visual language multimodal large model, crops the corresponding region from the original image and returns the local image. The local image returned by the tool is converted into a visual token and added back to the current context, so that the visual language multimodal large model can combine the original image and the local image to jointly determine the defect type and location in the subsequent generation process.

5. The insulator defect detection method based on multimodal large model tool invocation according to claim 1, characterized in that, Step S4 specifically includes: S4.1 Markov Decision Process Modeling: The detection process is modeled as a Markov decision process. Let the state at step t consist of the original image, detection instructions, historical generated text, and historical tool observations. In the current state, the model selects an action according to the strategy. The action is a plain text token or a tool call instruction. S4.2 Comprehensive Reward Function Design: The output format, category judgment, bounding box localization, and tool invocation effectiveness are all included in the reward. The total reward consists of four parts: format reward, classification reward, bounding box localization reward, and tool invocation reward. The format reward is used to constrain the correctness of the JSON format of the model output. The classification reward is used to evaluate the consistency between the model's predicted insulator state category and the true category. The bounding box localization reward is used to evaluate the spatial matching degree between the predicted bounding box and the true labeled box. The tool invocation reward is used to evaluate the effectiveness of the model in calling visual tools.

6. The insulator defect detection method based on multimodal large model tool calling according to claim 5, characterized in that: Bounding box localization reward in S4.2 Specifically, it includes: Intersection over Union (IoU) is used to measure the degree of overlap between predicted and ground truth bounding boxes. When the IoU is greater than a set threshold, the predicted bounding box is considered to be successfully matched with the ground truth bounding box. For cases with multiple defects in a single image, a greedy matching method is used to match the set of predicted bounding boxes with the set of ground truth bounding boxes. The number of successfully matched bounding boxes is counted, and the localization accuracy, recall rate, and average IoU are calculated. For normal samples, if the model outputs the category as "normal" and the bounding box set is empty, the localization result is considered correct. If the output is a non-empty bounding box, it is considered a false detection.

7. The insulator defect detection method based on multimodal large model tool calling according to claim 5, characterized in that: The tool call reward in S4.2 adopts a conditional reward strategy, specifically including: The conditional tool reward is calculated by subtracting the invalid call penalty from the valid call reward. The valid call reward is triggered only when both the classification reward and the bounding box reward are positive. The invalid call penalty is proportional to the number of invalid tool calls. A tiered reward system is used during training: a negative reward is given for format errors, a lower positive reward is given for correct category but insufficient localization, and a higher reward is given when both the category and the bounding box are correct.

8. The insulator defect detection method based on multimodal large model tool invocation according to claim 5, characterized in that, S4 also includes: S4.3, PPO Policy Optimization: An actor-critic architecture is used for training. A LoRA-tuned visual-language multimodal large model is used as the initial actor model, and a visual-language multimodal large model of the same scale is used as the critic model for state value estimation. The ratio of action probabilities between the old policy and the current policy is set as the probability ratio. The PPO pruning policy objective is the expectation of the smaller of the product of the probability ratio and the advantage function and the product of the pruned probability ratio and the advantage function. The advantage function is jointly determined by the reward signal and the state value estimated by the critic. During training, the policy gradient is calculated only for the text tokens generated by the model and the tool call actions. The visual observations returned by the tool are processed by loss masking.

9. The insulator defect detection method based on multimodal large model tool invocation according to claim 1, characterized in that, Step S5 specifically includes: S5.1 Preliminary judgment of the whole image: After receiving the inspection image of the insulator to be inspected, the visual language multimodal large model first extracts the features of the whole image through the visual encoder and makes a preliminary judgment. S5.2 Local Focus Verification: When the confidence level of the model in the preliminary judgment result is lower than the preset threshold, the visual language multimodal large model autonomous generation tool calls the instruction to output the coordinates of the candidate region, calls the image local cropping and magnification tool to locally magnify and observe the region, and the local image returned by the tool is added back to the context. The model combines the original image and the local image to perform verification judgment. S5.3 Structured Result Output: The visual language multimodal large model autonomously decides whether to make multiple rounds of tool calls based on the learned strategy until it obtains enough information to make a final judgment. The final output is kept in a unified JSON format. For defect samples, the output includes a list of defect types and bounding box coordinates. For normal samples, the output includes a list of normal types and empty bounding boxes.

10. An insulator defect detection system based on multimodal large model tool invocation, used to implement the method as described in any one of claims 1-9, characterized in that, include: The dataset building module is used to collect and preprocess images of substation insulator inspections, and generate multimodal training samples containing task prompts, images, and structured labels. The model loading and fine-tuning module is used to load pre-trained visual language multimodal large models and use the LoRA parameter efficient fine-tuning method to perform domain adaptation training on them. The tool integration module is used to embed visual processing tools into the inference process of a large multimodal visual language model, enabling tool call command parsing and multimodal context management; The reinforcement learning optimization module is used to optimize the defect detection strategy and tool calling behavior of the visual language multimodal large model by combining the PPO algorithm with a multi-dimensional comprehensive reward function. The inference detection module receives the image to be detected, drives the visual language multimodal large model to perform multi-round inference with tool calls, and outputs the structured defect detection results.